Successive Over Relaxation Q-Learning

In a discounted reward Markov Decision Process (MDP), the objective is to find the optimal value function, i.e., the value function corresponding to an optimal policy. This problem reduces to solving a functional equation known as the Bellman equation and a fixed point iteration scheme known as the...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	arXiv.org 2019-06
Hauptverfasser:	Kamanchi, Chandramouli, Raghuram Bharadwaj Diddigi, Bhatnagar, Shalabh
Format:	Artikel
Sprache:	eng
Schlagworte:	Algorithms Computer Science - Learning Convergence Fixed points (mathematics) Functional equations Iterative methods Machine learning Markov analysis Markov processes Mathematical models Statistics - Machine Learning
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Schreiben Sie den ersten Kommentar!