Distributionally Robust Warm Start for a Population of Bandits

Poster D: Tuesday -- 16:00 - 18:00

Sumantrak Mukherjee, Debabrota Basu, Jasmin Brandt, Viktor Bengs, Eyke Hüllermeier, Sebastian Josef Vollmer

Keywords: distributionally robust optimisation, multi-task bandits, warm-starting, meta-learning

Warm-starting a Bayesian multi-armed bandit algorithm for a population of bandits requires an initialization in the form of a prior. To this end, one usually makes use of previously collected data. However, in many real-world applications, the deployment distribution may differ from the distribution from which data was collected. In this work, we propose Distributionally Robust Warm-Starts (DRWS) for warm-starting under distribution shift, i.e., changes of the population or reward distributions. DRWS uses scalable regret proxies and provides guarantees relating proxy error, sample size, and uncertainty-set design to worst-case performance, as well as conditions for robustness gains over nominal warm-starts. Experiments on synthetic and semi-synthetic benchmarks show improved early-round and worst-case regret under realistic shifts.