BluffJAX: Adversarial Imperfect Information Games in JAX

Poster D: Tuesday -- 16:00 - 18:00

Aryaman Reddi, Jan Peters, Carlo D'Eramo

Keywords: benchmark, games, reinforcement learning, multi agent, multi agent reinforcement learning, marl, adversarial, game theory, game theoretic, poker, bluff, werewolf, exploitability, jax, GPU, vmap, jit, jit compile, jit compilation, parallelization

We introduce BluffJAX: an open-source suite of adversarial imperfect information games in JAX. We provide canonical implementations of games designed for high simulation throughputs and parallelization on GPU accelerators. Our suite consists of well-studied benchmarks such as Texas Hold’Em Poker and Kuhn Poker, as well as games that have not been studied in reinforcement learning research, such as Bluff, Stud Poker, and Kemps. We hope that implementing a variety of game mechanics and difficulties will introduce new challenges and foster novel research directions in game-theoretic methods for RL. We benchmark the throughput performance and memory usage of our environments in single and multi-GPU settings, demonstrating scaling of up to hundreds of millions of samples per second, and motivating the usage of BluffJAX over related GPU and CPU-based libraries. We benchmark reinforcement learning, tree search, and game-solving algorithms in JAX in order to provide users with baseline results and facilitate future comparisons.