Alterations in choice behavior by manipulations of world model

How to compute initially unknown reward values makes up one of the key problems in reinforcement learning theory, with two basic approaches being used. Model-free algorithms rely on the accumulation of substantial amounts of experience to compute the value of actions, whereas in model-based learning...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	Proceedings of the National Academy of Sciences - PNAS 2010-09, Vol.107 (37), p.16401-16406
Hauptverfasser:	Green, C. S., Benson, C., Kersten, D., Schrater, P.
Format:	Artikel
Sprache:	eng
Schlagworte:	Bayes Theorem Bayesian analysis Behavior Behavior modeling Biological Sciences Bullets Choice Behavior Decision making Ecological modeling Ecosystem Experimentation Female Humans Jet fighters Learning Machine learning Male Models, Biological Probabilities Probability Roulette Tunnels Young Adult
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	How to compute initially unknown reward values makes up one of the key problems in reinforcement learning theory, with two basic approaches being used. Model-free algorithms rely on the accumulation of substantial amounts of experience to compute the value of actions, whereas in model-based learning, the agent seeks to learn the generative process for outcomes from which the value of actions can be predicted. Here we show that (i) "probability matching"— a consistent example of suboptimal choice behavior seen in humans —occurs in an optimal Bayesian model-based learner using a max decision rule that is initialized with ecologically plausible, but incorrect beliefs about the generative process for outcomes and (ii) human behavior can be strongly and predictably altered by the presence of cues suggestive of various generative processes, despite statistically identical outcome generation. These results suggest human decision making is rational and model based and not consistent with model-free learning.
ISSN:	0027-8424 1091-6490
DOI:	10.1073/pnas.1001709107