Bellman-Admissible Scalar Losses: Dynamic Calibration, Fundamental Limits, and Misspecification Geometry

Poster B: Monday -- 16:00 - 18:00

Manoj Saravanan

Keywords: reinforcement learning, theoretical reinforcement learning, Bellman regression, value-based reinforcement learning, surrogate losses, Bellman calibration, Bregman divergence, proper scoring rules, misspecification geometry, finite-horizon MDPs

We study which surrogate losses are admissible for Bellman regression in finite-horizon reinforcement learning. We define \emph{Bellman-composability}: a loss and decoder are Bellman-composable if every Bayes-optimal fit to any Bellman target law decodes to its Bellman mean. In the scalar regime, we prove that Bellman-composability is equivalent to strict mean-consistency and hence, under standard regularity assumptions, to strict generalized Bregman mean scoring. We then establish a dynamic Bellman-calibration theorem: stagewise surrogate excess risk controls one-step Bellman residuals and therefore propagates to sup-norm \(Q\)-error and greedy-policy suboptimality. In the scalar Bregman regime, we show that excess surrogate risk is exactly a Bregman divergence to the Bellman mean, yielding exact calibration formulas and a universal scalar barrier: every smooth scalar Bellman-composable loss has worst-case calibration modulus of order \(\Theta(\sqrt{\varepsilon})\), and this order is witnessed by a one-step control lower bound. Finally, we prove a local misspecification-geometry theorem showing that admissible losses induce curvature-weighted projection operators; consequently, losses that agree on the target and on worst-case calibration can nevertheless exhibit different first-order misspecification amplification. Together, these results separate admissibility, calibration, and approximation geometry, and provide a theory-first foundation for loss design in value-based reinforcement learning.