Online Performative Decision Making with Latent Distribution States

Many operational decisions reshape the populations they act on: routing policies alter traffic, care interventions affect health outcomes, and public programs change participation. We study online control of such decision-dependent populations when the primitive state is a distribution, actions determine both current reward and the next distribution, and the reward and transition laws are unknown. Existing performative-prediction models largely ask whether repeated deployment converges to a stable decision-distribution pair. In many operational settings, however, the goal is sequential control of one evolving population, which cannot be reset for offline exploration. Exploration is therefore risky: an informative action can move the population to a state from which high reward is difficult to recover. We propose a latent-state formulation in which a finite-dimensional statistic summarizes decision-relevant information in the population, is estimated from finite batches, and evolves under unknown linear-in-feature reward and transition models. The key structural condition is recoverability: actions taken from one latent state can be matched at nearby latent states with comparable rewards and contracted next-state discrepancies. This converts transition-estimation error into controlled value error in continuous latent dynamics. Under robust recovery of the confidence class, full-model optimistic planning achieves high-probability sublinear regret. When only the true dynamics are recoverable, we introduce recoverable-core extended optimism, which optimizes over locally plausible reward-transition triples while pruning nonrecoverable ones. We give a finite-grid implementation using a Lipschitz majorant. The regret bounds separate reward learning, transition learning, epoch switching, finite-batch state estimation, and finite-grid approximation. A lower bound shows that without recovery-type structure, irreversible latent dynamics can force linear regret for every single-trajectory algorithm.

Article

Download

View PDF