Complexity Estimation for Q-Functions in Reinforcement Learning
Poster A: Monday -- 11:00 - 12:30
Nigel De Meulder, Ali Anwar, Siegfried Mercelis
Keywords: Reinforcement learning, Value function geometry, Perturbation sensitivity, Finite-difference methods, Stencil-based estimators
We study how learned value functions respond to input perturbations in reinforcement learning. We argue that sensitivity is inherently multi-scale and objective-dependent: worst-case changes are governed by first-order directional structure in the infinitesimal limit but require finite-scale treatment as perturbation size grows, while the magnitude of variation depends on both first- and higher-order effects and the aggregation of variation across directions. To capture these regimes, we introduce stencil-based estimators, building on classical finite-difference methods, that probe value functions at a finite spatial scale using only function evaluations. Empirically, finite-scale first-order estimators outperform infinitesimal gradients in predicting worst-case value drops at later stages of training, while a multi-order complexity measure provides the strongest signal for predicting the magnitude of variation under perturbations. These results highlight that the effectiveness of geometric descriptors depends on both the perturbation scale and the underlying sensitivity objective, and emphasize the importance of finite-scale analysis for understanding sensitivity in learned value functions.