Dyna-Style Safety Augmented Reinforcement Learning: Staying Safe in the Face of Uncertainty

Poster B: Monday -- 16:00 - 18:00

Artur Eisele, Bernd Frauenknecht, Friedrich Solowjow, Sebastian Trimpe

Keywords: Safe Exploration, Safe Reinforcement Learning, Model-Based Reinforcement Learning

Safety is a key problem in reinforcement learning, especially during exploration. While safety filters are promising to address this issue, they are ill-suited for high-dimensional systems with unknown dynamics. We propose Dyna-style Safety Augmented Reinforcement Learning (Dyna-SAuR), a novel algorithm that learns a control policy as well as a scalable safety filter requiring minimal domain knowledge by leveraging a learned uncertainty-aware dynamics model. A novel filter learning formulation enforces to avoid failures and areas of high model uncertainty. As joint deployment of the control policy and filter allow to safely collect data and improve the dynamics model, certain areas grow and filter conservatism decreases over time. Dyna-SAuR reduces failures compared to state-of-the-art methods by two orders of magnitude on goal-reaching CartPole and MuJoCo Walker.