Rationality Randomization with Maximum Entropy for Robust Dynamic Obstacle Avoidance on Legged Robots
Poster B: Monday -- 16:00 - 18:00
Jose-Luis Holgado-Alvarez, Gabriele Tiboni, Aryaman Reddi, Carlo D'Eramo
Keywords: Dynamic Obstacle Avoidance, Adversarial Reinforcement Learning, Maximum Entropy Curriculum, Legged Robot Navigation
Reinforcement learning (RL) for legged robots has recently gained significant attention due to its ability to enable traversal of complex terrains and navigation in cluttered environments. However, most existing approaches focus on static or predictably moving obstacles and rely on handcrafted curricula or manual task design, which may limit robustness to diverse obstacle behaviors.
In this paper, we propose a method to automatically generate rich and dynamic obstacle-avoidance scenarios for learning robust navigation policies. Inspired by adversarial reinforcement learning, we model obstacles as agents that induce challenging interactions for a protagonist policy. To improve robustness, we vary obstacle behavior by modulating their \textit{rationality}, defined as the level of stochasticity in their actions. This is controlled via a temperature parameter that interpolates between deterministic, goal-directed motion and random behavior.
We implement this through a domain randomization framework that samples rationality coefficients from a probability distribution, exposing the policy to a diverse range of obstacle behaviors during training. Crucially, this distribution is adapted online via a maximum entropy objective, which expands behavioral diversity while maintaining task performance.
We demonstrate that the proposed method, Rationality Randomization with Maximum Entropy (RAMEN), enables the learning of navigation policies that generalize across a wide range of simulated dynamic environments. Furthermore, we show successful zero-shot deployment on a legged robot operating in dynamic real-world scenarios using only onboard sensory inputs.