Approximate solution of infinite-horizon risk-sensitive Markov decision processes

Infinite-horizon risk-sensitive Markov decision processes (MDPs) under the discounted cost criterion are challenging to solve because the optimal policy may be nonstationary. Existing methodsĀ  typically reformulate the problem as a continuous-state risk-neutral MDP and rely on state discretization or value-function approximation, often without explicit stopping conditions or error bounds. In this paper, we present approximate variants of value iteration, linear programming (LP), and policy iteration to directly solve risk-sensitive MDPs without reformulation. We show that the approximate value functions obtained by value iteration and LP asymptotically converge to the optimal value function, and the policies generated by policy iteration converge in value to an optimal policy. We derive stopping conditions to achieve arbitrarily close approximations to the optimal value function. Our methods yield near-optimal policies that are nonstationary in only the first N periods and stationary thereafter, where N grows logarithmically in the inverse error tolerance.

Article

Download

View PDF