Inverse reinforcement learning in contextual MDPs

We consider the task of Inverse Reinforcement Learning in Contextual Markov Decision Processes (MDPs). In this setting, contexts, which define the reward and transition kernel, are sampled from a distribution. In addition, although the reward is a function of the context, it is not provided to the a...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	Machine learning 2021-09, Vol.110 (9), p.2295-2334
Hauptverfasser:	Belogolovsky, Stav, Korsunsky, Philip, Mannor, Shie, Tessler, Chen, Zahavy, Tom
Format:	Artikel
Sprache:	eng
Schlagworte:	Algorithms Artificial Intelligence Computational geometry Computer Science Control Convex analysis Convexity Decision making Empirical analysis Learning Machine Learning Markov processes Mechatronics Mortality Natural Language Processing (NLP) Optimization Patients Physicians Precision medicine Robotics Sepsis Simulation and Modeling Special Issue on Reinforcement Learning for Real Life
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	We consider the task of Inverse Reinforcement Learning in Contextual Markov Decision Processes (MDPs). In this setting, contexts, which define the reward and transition kernel, are sampled from a distribution. In addition, although the reward is a function of the context, it is not provided to the agent. Instead, the agent observes demonstrations from an optimal policy. The goal is to learn the reward mapping, such that the agent will act optimally even when encountering previously unseen contexts, also known as zero-shot transfer. We formulate this problem as a non-differential convex optimization problem and propose a novel algorithm to compute its subgradients. Based on this scheme, we analyze several methods both theoretically, where we compare the sample complexity and scalability, and empirically. Most importantly, we show both theoretically and empirically that our algorithms perform zero-shot transfer (generalize to new and unseen contexts). Specifically, we present empirical experiments in a dynamic treatment regime, where the goal is to learn a reward function which explains the behavior of expert physicians based on recorded data of them treating patients diagnosed with sepsis.
ISSN:	0885-6125 1573-0565
DOI:	10.1007/s10994-021-05984-x