Differentiating Through Moving Recourse: Feasible Policy Optimization and Finite-Sample Certification for Multistage Stochastic Programs
Multistage stochastic programs model decisions under uncertainty where earlier decisions change later feasible sets. We propose a feasible pathwise gradient method (FeasPG) that can update the policy as new sample paths arrive. The policy proposes a decision and projects it onto the current feasible set, so every decision is feasible. We derive the full trajectory … Read more