Co-Exploration and Co-Exploitation via Shared Structure in Multi-Task Bandits
Poster C: Tuesday -- 11:00 - 12:30
Sumantrak Mukherjee, Serafima Lebedeva, Jasmin Brandt, Valentin Margraf, Jonas Hanselle, Kanta Yamaoka, Viktor Bengs, Stefan Konigorski, Eyke Hüllermeier, Sebastian Josef Vollmer
Keywords: Multi-armed bandits, Multi-task bandits, meta learning, non-parametric, gaussian processes, Thompson sampling
We propose CoCo-Bandits, a Bayesian framework for efficient exploration in contextual multi-task multi-armed bandits with partially observed context and latent reward dependencies across tasks and arms. Our approach learns a global joint distribution across tasks while retaining personalised inference, and distinguishes structural uncertainty from user-specific uncertainty due to incomplete context and limited history. Concretely, we approximate the joint distribution over tasks and rewards with a particle-based log-density Gaussian Process, enabling discovery of inter-arm and inter-task dependencies without parametric latent assumptions. We show theoretically that learning the meta-prior yields sublinear cumulative multi-task regret under posterior contraction assumptions. Empirically, on synthetic benchmarks and a real-world-dataset-based environment, our method outperforms parametric baselines and shows that nonparametric latent modelling improves robustness under misspecification and complex heterogeneity.