Who's Winning? Identifying Nash Equilibrium from Improvement Feedback

Poster C: Tuesday -- 11:00 - 12:30

Cyrille Kone, Giorgia Ramponi

Keywords: multi-agent reinforcement learning, pure exploration, markov games, adpative stopping, preference

We study pure exploration in two-player zero-sum discounted Markov games under preference feedback. The agents never observe rewards. Instead, they receive only intermittent binary comparisons over trajectory segments, generated from an unknown latent reward model through a Bradley--Terry mechanism. The goal is to identify an $\varepsilon$-Nash equilibrium with high probability while minimizing the expected stopping time. We adopt a pure exploration perspective in which each agent seeks to terminate the game as quickly as possible while guaranteeing that, at stopping, its recommended policy constitutes a dominant strategy. Stopping is negotiated in a fully decentralized manner: each agent may independently emit a stopping flag, and the game terminates as soon as both agents have done so, at which point each player commits to its recommended strategy.