Times Series Meet MDPs for Patient Follow Up
Poster C: Tuesday -- 11:00 - 12:30
Michalak Nicolas, Emilie Kaufmann, Timothée Mathieu, Philippe Preux
This paper introduces a novel reinforcement learning problem motivated by automated patient follow-up. Trajectories of different patients are monitored in parallel and medical interventions should be selected adaptively so as to stabilize the patients' trajectories while minimizing the total cost of the interventions.
In this setting, each patient generates a unique trajectory with unknown, trajectory-specific parameters, which makes classical reinforcement learning applied to each patient ineffective. We propose a controlled time-series model for the evolution of each patient, that can then be modeled by a parametric Markov Decision Process. These different MDPs depend on an individual patient parameter and some global intervention parameters, that should be learnt jointly in order to adopt a near-optimal follow-up strategy. We propose a Thompson Sampling follow-up algorithm in which the posterior sampling step is approximated using Gibbs sampling. We perform a thorough empirical study revealing the inefficiency of classical RL algorithms and the good performance of our Thompson Sampling strategy.