Who's Winning? Identifying Nash Equilibrium from Improvement Feedback
Poster C: Tuesday -- 11:00 - 12:30
Cyrille Kone, Giorgia Ramponi
Keywords: multi-agent reinforcement learning, pure exploration, markov games, adpative stopping, preference
We study pure exploration in two-player zero-sum discounted Markov games under preference feedback. The agents never observe rewards. Instead, they receive only intermittent binary comparisons over trajectory segments, generated from an unknown latent reward model through a Bradley--Terry mechanism. The goal is to identify an $\varepsilon$-Nash equilibrium with high probability while minimizing the expected stopping time. We adopt a pure exploration perspective in which each agent seeks to terminate the game as quickly as possible while guaranteeing that, at stopping, its recommended policy constitutes a dominant strategy. Stopping is negotiated in a fully decentralized manner: each agent may independently emit a stopping flag, and the game terminates as soon as both agents have done so, at which point each player commits to its recommended strategy.