Multistage stochastic programs model decisions under uncertainty where earlier decisions change later feasible sets. We propose a feasible pathwise gradient method (FeasPG) that can update the policy as new sample paths arrive. The policy proposes a decision and projects it onto the current feasible set, so every decision is feasible. We derive the full trajectory gradient, accounting for how earlier decisions change later constraint bounds. Under regularity conditions, it is unbiased at almost every parameter value. We also bound the error from treating the constraint bounds as fixed during differentiation. We analyze a proposal-distance penalty that can provide a training signal where projection makes cost locally flat. We give a decreasing-weight schedule that preserves the O(1/√K) stationarity rate for the expected cost over K training updates under standard stochastic-gradient assumptions. To assess the learned policy, we use held-out trajectories to construct a finite-sample upper confidence bound on its optimality gap. The certificates are also anytime-valid across checkpoints, so the bounds can guide when training stops. In reservoir-management experiments, FeasPG obtains lower mean costs than stochastic dual dynamic programming and reinforcement-learning baselines on model-generated paths. A separate reservoir study demonstrates informative certificates with no coverage failures in 200 repetitions per setting.