Reinforcement Learning with Complex (valued) Memories

Poster B: Monday -- 16:00 - 18:00

Sathya Kamesh Bhethanabhotla, Stratis Gavves, André Biedenkapp

Keywords: Reinforcement Learning, Complex-Valued Neural Networks, Partial Observability, Unitary Recurrent Networks, Quantum Information Theory

Partially observable environments pose a fundamental challenge in deep reinforcement learning, requiring agents to compress temporal information from observations and maintain a memory to make effective decisions. While there exist many approaches ranging from gated recurrence to attention mechanisms and model-based RL, the search for effective representational techniques that can capture long-term dependencies remains an active area of research. In this work we revisit Unitary recurrent networks (uRNNs) [Arjovsky et al., 2016, Jing et al., 2017], that demonstrated superior gradient flow and associative recall, expressing the recurrence and the hidden state in a complex vector space. Their norm preserving unitary dynamics enable information propagation through long sequences. To this end, we propose three different versions of uRNNs as drop-in replacements for recurrent PPO architectures, and demonstrate that the simple recurrence and the added degree of freedom from the phase of the complex representations enable significant gains over baselines on several memory-improvable tasks, including continuous control. We further explore how to preserve the phase information of the complex hidden state for a \textit{phase-aware} policy by drawing a parallel to how quantum states are measured. With our methods reaching up to $2$-$3\times$ the reward in environments like rocksample and Craftax compared to the baselines, this work points towards an exciting new direction of representations for RL and the problem of partial observability.