Inferring Transferable Rewards via Active Inverse Reinforcement Learning

Poster B: Monday -- 16:00 - 18:00

Victor Villin, Till Freihaut, Andreas Schlaginhaufen, Maryam Kamgarpour, Christos Dimitrakakis, Giorgia Ramponi

Keywords: inverse reinforcement learning, active learning, environment design, transferability, identifiability, robustness, alignment

Inverse Reinforcement Learning (IRL) aims to recover expert rewards that transfer across environments with differing dynamics. Prior work has shown that such transfer is possible when gathering expert demonstrations across sufficiently different dynamics. In practice, however, expert demonstrations are costly, making it critical to select the right dynamics to reduce the quantity of expert data required for learning transferable rewards. We propose a two-phase approach for doing so: we cycle through the available dynamics to obtain reliable initial reward estimates, before we actively choose them by minimizing uncertainty about the expert reward. Crucially, this leads to new sample complexity guarantees in terms of transferability. Empirically, our method significantly reduces the amount of expert data needed to learn a transferable reward.