Liyu Xia, Sarah L. Master, Maria K. Eckstein, Beth Baribault, Ronald E. Dahl, Linda Wilbrecht and Anne Gabrielle Eva Collins. Modeling changes in probabilistic reinforcement learning during adolescence. PLOS Computational Biology 17(7):e1008524. Published article, 1 July 2021. Full reading completed 29 September 2026.
Reading coverage and retained sources
Read all 22 main PDF pages, including methods, equations, declarations, references and supporting-information captions; all seven pages of S1 Text; and all 12 supplementary figures and four supplementary tables. Visually inspected main Figures 1–5 and the equations on pages 7–8. All 16 supplementary TIFFs contain two frames; each pair is pixel-identical (ImageMagick absolute-error count zero), so one frame of every pair was inspected. This is a completed reading, not an abstract-based account.
Retained main PDF,
main text,
S1 Text, and supplementary
TIFFs s002–s017 under the same filename prefix. The
provenance manifest
records all source URLs, sizes, hashes and exact coverage. Sources are CC BY 4.0.
No model was reproduced and no cited paper is included in this reading. Public data
and code audits are separate work, not evidence supplied by reading this article.
Question, cohort and task
The study asks how performance and fitted learning processes vary across ages 8–30 in a stable probabilistic environment. This is a cross-sectional developmental study, not observation of the same person's development. Of 297 completers, 187 were youth, 55 university students and 55 community adults. One person with only 18 trials was removed; 21 were removed for switching more than staying after positive feedback for the same stimulus; another 11 met both a poor-performance condition and a behavioral or missingness condition. The final sample is 264, including 138 females and 157 people younger than 18. The authors report qualitatively similar principal findings with weaker exclusions, but those checks do not undo selection in the public analytic data.
Participants saw one of four butterflies and selected one of two flowers. Each butterfly preferred one flower with 80% reward probability, versus 20% for the other. Associations were stable. The task had 120 feedback trials, 30 per butterfly; trial order and outcomes were predetermined with counterbalancing/randomization. Choice deadline was seven seconds, followed by one second showing the selection and two seconds of feedback. Earned points did not determine an additional monetary bonus. It was the third of four tasks during the visit. The paper alone does not supply a cross-release participant-ID mapping or authenticate every released record's version.
S1 Text gives the final exclusion rule precisely: at/below-chance performance and one of extreme overall staying, extreme switching, more than 12 consecutive stays, or incomplete data. Table S1 shows eight of these 11 exclusions among ages 8–13, two among 13–18 and one among 25–30. This matters for developmental generalization and for any new prediction study that must avoid defining eligibility using future choices.
Behavioral analysis and computational models
Analyses examine accuracy, reaction time, feedback sensitivity, accumulated rewards for the currently presented stimulus and delay since its previous reward. More prior rewards and shorter delay predict better performance; sensitivity to accumulated reward becomes stronger with age. These descriptive analyses use experimenter-known correctness as an outcome. They do not license supplying hidden correctness to a behavioral predictor.
All Q-learning candidates start action values at 0.5 and choose with a softmax whose inverse temperature is β. The selected stimulus/action value moves toward binary feedback, with a single rate α or separate positive/negative prediction-error rates α+ and α−. Some models fix α− to zero. Forgetting moves every value except the currently chosen stimulus/action pair toward 0.5 on each trial. A model with this forgetting operation therefore differs from one that simply decays the previous stimulus's memory between its presentations.
The six candidates combine rate asymmetry, zero negative learning and forgetting. Hierarchical fits use truncated-normal person distributions and weak uniform hyperpriors, with four NUTS chains, 4,000 retained iterations per chain after 1,000 warmup iterations. The αβf model did not converge hierarchically and was fit independently for its reported simulations. Comparisons involving that model therefore do not all share one estimation procedure.
The full α+α−βf model has the lowest WAIC. Figure 2's differences relative to it are 1,705 for αβ, 717 for α+0β, 509 for α+α−β and 123 for α+0βf. Nevertheless the authors select α+0βf for interpretation: fitted α− was close to zero and poorly recovered, and full-model simulations overshot observed learning curves. The selected model tracks those curves better. This is a combination of fit, recoverability and posterior-predictive judgment, not a held-out chronological model-selection result.
The supplementary checks materially qualify that choice. With α− generated over a broader, higher range, it becomes recoverable (Figure S3). Thus failure to recover near-zero α− does not show that negative learning is fundamentally unidentifiable. Likewise setting α− to zero is a modeling decision, not evidence that humans never learn from unfavorable feedback. Parameter simulations also show that modest positive α− can perform better than zero in the task, even when α− below α+ is often advantageous.
The paper generates 100 posterior-predictive simulations per person and checks learning curves and recovered parameters/age coefficients. Its separate model-confusion check uses only three independently fit models and one simulated dataset per generating model. Table S3 gives rounded protected exceedance probabilities of one on its diagonal and zero off it. This is not a complete recovery test of all six hierarchical procedures; the supplement explicitly cites their computational cost as the reason for reducing this check. Figure S6's flat-fit BIC comparison favors α+0βf across age groups, a useful consistency check with that different procedure.
Findings, supplementary evidence and limits
Performance improves through adolescence and approaches a plateau around age 19; reaction times decrease. The reported pattern is not a mid-adolescent performance peak in this stable task. In the selected model, positive learning rate and inverse temperature increase with age with curvature; forgetting has weaker evidence for an age association. The very simple symmetric model yields much smaller learning rates, illustrating that a parameter's magnitude depends on which other processes the model can express. It cannot be compared across specifications as an interchangeable trait.
Figures S7–S9 show nonmonotonic, interacting performance surfaces for α+, α−, β and forgetting. Optimal values depend on the other parameters; eliminating forgetting is not uniformly optimal. These are simulations under specific task/model assumptions, not estimates of an intrinsic human optimum. Figure S5 supports recovery of fitted age effects under the generating assumptions. Figure S10 reports age effects with a simpler model; consistency of some trends does not identify one unique mechanism.
Pubertal-development scores, testosterone and age are correlated (Figure S11). Most puberty associations weaken or disappear after accounting for age. An association between testosterone and α+ in ages 13–15 survives the stated multiple-comparison adjustment; the authors explicitly allow that this exploratory result may be a false positive. Figure S12 and Tables S2/S4 expose small subgroup sizes, imperfectly matched age ranges and within-bin variance/correlation. No causal hormonal conclusion follows.
Other limits include different recruitment and compensation across older age cohorts, performance-based exclusions that disproportionately remove younger people, a single task and visit, and the distinction between age differences and within-person change. The task's lack of performance-contingent payment may also affect engagement. Model simulation agreement on the same data is neither new-person generalization nor prediction from a person's earlier task into a later one.
Consequences for the current personal-learning research
This paper provides a well-specified earlier stable-learning context, with multiple stimuli and uncertain feedback, for the candidate chronological A→B study. It supports allowing context to change the way personal information is expressed. It does not directly test whether earlier personal data improves a population learner's later-choice forecast, or whether personal weights add to a shared model given the same history.
An appropriate population reference should have access to available age/cohort context; otherwise personal history can appear useful merely by revealing population structure. Earlier personal learning may still add useful information after that control. Its contribution could be beliefs, memory, attention, exploration or response behavior; we should not reduce the question to a universal fitted learning rate.
The original article's 80/20 schedule conflicts with the context paper's 70/30 description in one section (that paper elsewhere says 80/20). Our separate code audit finds 80/20 acquisition code and an additional 32-trial no-feedback phase. Neither observation alone proves the provenance of a particular released table. In particular, zero reward during a no-feedback phase must not be treated as observed negative feedback. Exact boundaries, coding and identity linkage belong in the dataset audit.
As of this reading, the public A table uses contiguous subject labels 1–264; these cannot safely be joined by numeric equality to the original IDs in the B release. See the separate code audit. Do not fit cross-task personal models until a documented linkage is established. This is a data-feasibility constraint, not a negative finding about personal learning.