Inferring Transferable Rewards via Active Inverse Reinforcement Learning
Poster B: Monday -- 16:00 - 18:00
Victor Villin, Till Freihaut, Andreas Schlaginhaufen, Maryam Kamgarpour, Christos Dimitrakakis, Giorgia Ramponi
Keywords: inverse reinforcement learning, active learning, environment design, transferability, identifiability, robustness, alignment
Inverse Reinforcement Learning (IRL) aims to recover expert rewards that transfer across environments with differing dynamics. Prior work has shown that such transfer is possible when gathering expert demonstrations across sufficiently different dynamics. In practice, however, expert demonstrations are costly, making it critical to select the right dynamics to reduce the quantity of expert data required for learning transferable rewards.
We propose a two-phase approach for doing so: we cycle through the available dynamics to obtain reliable initial reward estimates, before we actively choose them by minimizing uncertainty about the expert reward. Crucially, this leads to new sample complexity guarantees in terms of transferability. Empirically, our method significantly reduces the amount of expert data needed to learn a transferable reward.