A Simple Scaling Model for Bootstrapped DQN

Poster E: Wednesday -- 11:00 - 12:30

Roman Knyazhitskiy, Pascal R. Van der Vaart

Keywords: Reinforcement Learning, Epistemic Exploration

We present a large-scale empirical study of Bootstrapped DQN (BDQN) and Randomized-Prior BDQN (RP-BDQN) in the DeepSea environment designed to isolate and parameterize exploration difficulty. Our primary contribution is a simple scaling model that accurately captures the probability of reward discovery as a function of task hardness and ensemble size. This model is parameterized by a method-dependent effectiveness factor, $\psi$. Under this framework, RP-BDQN demonstrates substantially higher effectiveness ($\psi \approx 0.87$) compared to BDQN ($\psi \approx 0.80$), enabling it to solve more challenging tasks. Our analysis reveals that this advantage stems from RP-BDQN's sustained ensemble diversity, which mitigates the posterior collapse observed in BDQN. Furthermore, we show how systematic deviations from this simple model diagnose complex second-order dynamics: we mechanistically link "cooperative" deviations in small ensembles to information sharing via the replay buffer, and "saturation" in large ensembles to correlated failures. We conclude by translating these findings into a prescriptive framework, offering practical guidance for configuring ensembles in deep exploration.