The Surprising Effectiveness of Approximate Value Iteration in Self-Play
Poster B: Monday -- 16:00 - 18:00
Raphael Boige, Amine Boumaza, Bruno Scherrer
Keywords: games, reinforcement learning, alphazero, planning, dynamic programming, value iteration, connect four, minimax
Combining search with function approximation has driven recent breakthroughs in game-playing programs, making self-play algorithms more competitive than ever. Still, the computational overhead of the most popular methods, based on MCTS, can be excessive. In this work, we investigate whether simpler methods remain competitive in non-trivial games of moderate-size like Connect4, Hex(7x7) and synthetic games. We evaluate a minimal self-play implementation of Approximate Value Iteration (AVI) using a ground-truth oracle for exact evaluation. Contrary to expectations, our results demonstrate the surprising effectiveness of AVI: it learns more accurate value functions and policies than AlphaZero, while being computationally lighter and practically simpler to implement. These findings suggest that the dominance of MCTS-based methods may have eclipsed simpler approaches that are competitive given modern hardware and software capabilities.