Approximate solution of infinite-horizon risk-sensitive Markov decision processes
Infinite-horizon risk-sensitive Markov decision processes (MDPs) under the discounted cost criterion are challenging to solve because the optimal policy may be non- stationary. Existing solution methods reformulate the problem as a continuous-state (risk-neutral) MDP and solve it using state-discretization or value function approximation. Such approaches typically lack explicit stopping conditions or error bounds. In this … Read more