Exploring the Role of Dopamine in Curiosity-Driven Learning Using Reinforcement Learning Models

Authors

  • Meixin Zhou

DOI:

https://doi.org/10.61173/970gv177

Keywords:

Curiosity-Driven Learning, Dopamine, Reinforcement Learning, Intrinsic Motivation

Abstract

Curiosity-driven learning is a core feature of biological intelligence, manifested in the active exploration of environments by organisms even in the absence of external rewards. This study integrates neurobiological findings with reinforcement learning theory to construct a unified computational framework that reconceptualizes dopamine as an encoder of "total prediction error." In this framework, behavioral selection maximizes total value, which comprises both external rewards and intrinsic informational value. Intrinsic rewards are quantified based on novelty or prediction error, while the weight assigned to exploration is dynamically modulated by the prefrontal cortex. Dopamine neurons encode total prediction error, thereby extending the conventional role of dopamine from a reward prediction error encoder to a more general encoder of violated expectations. The hippocampus generates intrinsic reward signals, the prefrontal cortex regulates the exploration weight, the ventral tegmental area integrates these signals to compute total prediction error, and the striatum translates this error into action selection. This model bridges the explanatory gap between neuroscience and computational theory, offering a unified perspective on the intrinsic motivation underlying curiosity and exploratory behavior. In practical terms, it also provides new insights into the mechanisms of psychiatric disorders.

References

[1] Schultz, W. (2019). Recent advances in understanding the role of phasic dopamine activity. F1000Research, 8, F1000- Faculty.

[2] Spanagel, R., & Weiss, F. (1999). The dopamine hypothesis of reward: past and current status. Trends in neurosciences, 22(11), 521-527.

[3] Part, A. (2000). CURRICULUM VITAE (CVA). Computer Science, 79, 164.

[4] LIU Huimin. (2024). Exploration situation and development strategy of Shengli Oilfield during“ 14th Five-Year Plan in 2021-2025 ”[J]. Petroleum Geology and Recovery Efficiency, 31(4):1~12

[5] Chirimuuta, M. (2014). Minimal models and canonical neural computations: The distinctness of computational explanation in neuroscience. Synthese, 191(2), 127-153.

[6] Barto, A. G. (2012). Intrinsic motivation and reinforcement learning. In Intrinsically motivated learning in natural and artificial systems (pp. 17-47). Berlin, Heidelberg: Springer Berlin Heidelberg.

[7] Zhu, Y., & Zhu, S. C. (2026). Utility. In Computer Vision: Cognitive Models for Visual Commonsense (pp. 213-242). Cham: Springer Nature Switzerland.

[8] Perez, S. M., & Lodge, D. J. (2018). Convergent inputs from the hippocampus and thalamus to the nucleus accumbens regulate dopamine neuron activity. Journal of Neuroscience, 38(50), 10607-10618.

[9] Targa Dias Anastacio, H., Matosin, N., & Ooi, L. (2022). Neuronal hyperexcitability in Alzheimer’s disease: what are the drivers behind this aberrant phenotype?. Translational psychiatry, 12(1), 257.

[10] Salaka, R. J., Nair, K. P., Annamalai, K., Srikumar, B. N., Kutty, B. M., & Shankaranarayana Rao, B. S. (2021). Enriched environment ameliorates chronic temporal lobe epilepsy‐induced behavioral hyperexcitability and restores synaptic plasticity in CA3–CA1 synapses in male Wistar rats. Journal of Neuroscience Research, 99(6), 1646-1665.

Downloads

Published

2026-06-24