On Natural Policy Compression

Poster A: Monday -- 11:00 - 12:30

Leonardo Cesani, Davide Tenedini, Matteo Papini, Marcello Restelli

Keywords: reward-free reinforcement learning, information geometry, natural policy gradient

Natural Policy Gradient is the theoretically principled approach to policy optimization: by following the geometry of the statistical manifold of policy distributions rather than the Euclidean geometry of the parameter space, it identifies the true direction of steepest ascent in behavior space. In practice, however, it is often computationally intractable, requiring online estimation and inversion of the Fisher Information Matrix at every iterate. In this paper, we pursue an alternative route: rather than fixing the optimizer, we ask whether the parameter space itself can be reparametrized offline, once and for all, so that standard Euclidean optimization becomes equivalent to natural gradient ascent. In the linear-Gaussian policy setting, we answer affirmatively. We construct an explicit Fisher-whitened embedding in which the squared Euclidean distance is proportional to the average KL divergence between policies, and show that subsequent PCA compression minimizes a Fisher-weighted reconstruction error, yielding a low-dimensional latent space in which vanilla gradient steps recover projected natural-gradient directions. We empirically validate the exactness of this construction in the linear-Gaussian setting. We then use this theoretical framework as a lens to interpret action-based neural policy compression, demonstrating through condition number analysis that its empirical success is explained by how faithfully its latent bottleneck approximates the geometric ideal identified by our analysis.