Robustness Is Free ? Global Convergence of Robust Policy Gradient Without Smoothing

Poster A: Monday -- 11:00 - 12:30

Navdeep Kumar, Kfir Yehuda Levy, Shie Mannor

Keywords: Robust reinforcement learning, Robust markov decision processes, Policy gradient methods, Non-smooth optimization, Sa-rectangular uncertainty sets

Robust policy gradient methods are a fundamental approach for robust reinforcement learning, but existing global convergence guarantees for general robust Markov decision processes (MDPs) rely on smoothing techniques such as Moreau envelopes due to the non-differentiability of the robust return. These approaches optimize smooth surrogate objectives and subsequently transfer guarantees back to the original robust problem, resulting in technically involved analyses and slow convergence rates. In this work, we show that smoothing the robust objective is unnecessary for establishing global convergence of robust policy gradient methods. Our analysis operates directly on the original non-smooth robust return by leveraging structural properties of MDPs. Specifically, we establish a smoothness-free sufficient increase guarantee under a weaker Lipschitz continuity assumption on the robust Q-function, and show that this property holds for general sa-rectangular uncertainty sets. We further develop a projection-geometry argument connecting sufficient increase and gradient domination guarantee in the non-smooth setting. Combining these ingredients, we establish a global convergence rate of $O(SAH^5\epsilon^{-1})$ for projected robust policy gradient methods in \texttt{sa}-rectangular robust MDPs, improving upon the previous $O(S^3A^5H^{14}\epsilon^{-4})$ guarantees obtained via smoothing techniques for general robust MDPs where $S,A,H$ is cardinality of state-space, action-space and horizon respectively. Notably, our complexity matches that of non-robust policy gradient methods up to lower-order factors, suggesting that robustness can be achieved without incurring substantial additional optimization cost.