Near-optimal Optimistic Reinforcement Learning using Empirical Bernstein Inequalities

We study model-based reinforcement learning in an unknown finite communicating Markov decision process. We propose a simple algorithm that leverages a variance based confidence interval. We show that the proposed algorithm, UCRL-V, achieves the optimal regret \(\tilde{\mathcal{O}}(\sqrt{DSAT})\) up...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	arXiv.org 2019-12
Hauptverfasser:	Tossou, Aristide, Basu, Debabrota, Dimitrakakis, Christos
Format:	Artikel
Sprache:	eng
Schlagworte:	Algorithms Communication Confidence intervals Decision theory Lower bounds Machine learning Markov analysis Markov chains
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	We study model-based reinforcement learning in an unknown finite communicating Markov decision process. We propose a simple algorithm that leverages a variance based confidence interval. We show that the proposed algorithm, UCRL-V, achieves the optimal regret \(\tilde{\mathcal{O}}(\sqrt{DSAT})\) up to logarithmic factors, and so our work closes a gap with the lower bound without additional assumptions on the MDP. We perform experiments in a variety of environments that validates the theoretical bounds as well as prove UCRL-V to be better than the state-of-the-art algorithms.
ISSN:	2331-8422