Learning Rewards in Strategic Games

Poster B: Monday -- 16:00 - 18:00

Antoine Bergerault, Cyrille Kone, Giorgia Ramponi

Keywords: inverse reinforcement learning, multi-agent, stackelberg games

Inverse reinforcement learning in multi-agent systems aims to recover latent reward functions from strategic interactions. Such problems naturally arise in applications including cybersecurity, autonomous driving, and online markets, where agents strategically adapt to one another while their underlying objectives remain unknown. While recent works have characterized feasible reward sets in games, it remains unclear how to extract rewards that preserve strategic behavior at deployment. In this work, we study this problem under bounded-rationality assumptions through the lens of strategic identifiability. We first show that learning from a fixed policy pair is fundamentally insufficient to guarantee low exploitability, even with infinite data. Motivated by this impossibility result, we propose an active inverse reinforcement learning framework for repeated Stackelberg games, where a leader adaptively selects policies to maximize information about the follower’s reward. Leveraging maximum likelihood estimation and experimental design, we derive finite-sample guarantees on the Stackelberg sub-optimality of the learned reward. We further extend our analysis to two-player zero-sum games, providing guarantees on preserving approximate Nash equilibria. Finally, we validate our approach on synthetic normal-form games and a wildlife protection simulation, showing the benefits of active exploration for reward recovery in strategic environments.