Control-centric Representation Learning using Action-free Datasets with Distinct Policies
Poster D: Tuesday -- 16:00 - 18:00
Max Rudolph, Rohan Patel, Alexander Levine, Peter Stone, Amy Zhang
Keywords: controllability, action-free learning, representation learning
Video, data without action or reward labels, can offer an abundant source of world knowledge in many tasks, such as robot manipulation and video game play. Many recent works have proposed methods to pre-train feature encoders from these data modalities in order to later fine-tune on less-abundant action-labeled, task-specific data. However, these methods fail when observations contain time-correlated noise. Critically, in many common visual control tasks, a naive encoder cannot distinguish noise from controllable features, leading to poor generalization. In this work, we consider a setting where video data is available from agents executing different policies operating in the same environment. We introduce a loss function that provably filters out time-correlated noise in the multi-policy video setting. We empirically demonstrate its capabilities on challenging, distraction-filled versions of DeepMind Control Suite, OGBench, and LIBERO tasks and can outperform prior action-free learning methods by 2$\times$. This work represents the first scalable method that can filter exogenous noise from high-dimensional, continuous observation and state spaces \textit{without} action labels.