Discovery of Sticky and Responsible Markov Options for Frozen-Option Transfer
Poster A: Monday -- 11:00 - 12:30
Yamen Habib, Dmytro Grytskyy, Rubén Moreno-Bote
Keywords: Option discovery, Frozen-option transfer, Transfer reinforcement learning
A fundamental problem in hierarchical reinforcement learning is finding a small number of reusable and easily composable behaviors that transfer with little or no fine-tuning to solve novel tasks. Option and skill discovery frameworks are two prominent approaches that address this problem. However, option discovery typically focuses on learning specialized behaviors for a training task, leading to a low diversity of behaviors with limited transferability, and unsupervised skill discovery learns diverse behaviors that can be misaligned with downstream tasks. To address these issues, we introduce Task-Agnostic Sticky and Responsible Options (TASRO), an option-learning framework where the high-level controller chooses low-level options using a fully option-switching Markovian structure, options are promoted to be persistent (sticky) and specialized (responsible), and the high-level controller has both task and task-agnostic information, while the low-level options only have task-agnostic information. We evaluate under a strict frozen-option transfer protocol: pretrain options on a goal-reaching training task, freeze them, and relearn only the controller on downstream tasks. Across navigation, locomotion, and manipulation domains, TASRO improves frozen-option transfer over baselines that rely on skill discovery or even when options are exposed with task information.