Maria K. Eckstein, Sarah L. Master, Ronald E. Dahl, Linda Wilbrecht and Anne G. E. Collins. “Reinforcement learning and Bayesian inference provide complementary models for the unique advantage of adolescents in stochastic reversal.” Developmental Cognitive Neuroscience 55, 101106 (2022), online 22 April 2022. DOI, PMC, author-hosted final PDF. The final article states CC BY-NC-ND 4.0. This summary is our paraphrase and assessment.
Read on 2026-09-29. Complete reading: the final 16-page main paper, including abstract, Methods, declarations, data availability and reference list; the entire 37-page scientific supplement, printed pages 67–103, §§6.1–6.4.3. Main Figures 1–6 and Tables 1–3, supplementary Figures 7–21 and Tables 4–17 were read and visually inspected. All 53 PDF pages were rendered and inspected; Figure 21 also received a separate enlarged inspection. Equations were read, not independently proved; plots were not digitized. PMC lists one unique scientific supplement, mmc1.pdf, twice in its XML. No cited paper was recursively opened and no model was run. Sources, extracted text, render hashes and exact reading scope are recorded in provenance.
Question, sample and task
The study asks whether learning in a volatile, stochastic environment improves continuously with age or instead shows an adolescent advantage, and whether reinforcement learning (RL) and Bayesian inference (BI) provide complementary explanations. It supplies positive evidence for structured, context-sensitive human learning and individual differences in fitted learning dynamics. It does not evaluate a model acquired from earlier personal experience on a later task or session.
The cross-sectional sample contains 291 people aged 8–30. Recruitment began with 191 community children/adolescents, 55 community adults and 66 university participants. Of the younger participants, 184 completed the task; five were excluded for mean accuracy below 58%, leaving 179. Two university participants were over 30 and seven lacked age information, leaving 57; all 55 community adults were retained. Community and university participants differed in recruitment, compensation and total testing duration. Participants completed four tasks during one visit. This paper does not give their exact order; order information must come from the separately audited cross-task report, not be attributed to this Methods section.
The task presents two boxes. One box yields a coin on 75% of choices, the other never does. The rewarding side reverses without a cue after a block-specific, prerandomized criterion of 7–15 collected rewards, allowing any number of unrewarded trials between them. Reversals occur only following a reward, and the first correct choice after a reversal is always rewarded while preserving the intended overall reward fraction. Participants must balance persistence after ambiguous nonreward against switching when contingencies have changed.
The tutorial comprises 10 deterministic trials followed by two deterministic reversal phases of eight trials each: 26 practice trials, then 120 main-task trials. Figure 1 specifies a choice deadline below five seconds, 0.2-second delay, one-second feedback and 0.5-second intertrial interval. The Methods explicitly say invalid trials were removed, leaving some participants fewer than 120 valid trials. Points and completed blocks are therefore rescaled as 120 × observed count / valid trials. The paper does not provide a raw-position mask or an exact missing-trial state-update convention. These omissions matter when reproducing the cleaned export but do not turn nonresponses into unrewarded choices.
Behavioral evidence and its strength
Participants typically begin favoring the new correct box roughly two trials after reversal and approach 80% accuracy after three to four trials. Positive feedback promotes staying more strongly than negative feedback promotes switching. Mixed-effects regressions distinguish rewarded and unrewarded choices at lags one through eight. Mid-adolescents show weaker responses to a single recent negative outcome but stronger integration of longer-range negative evidence, consistent with useful adaptation to ambiguous feedback.
Several performance measures peak around ages 13–15. The formal two-line tests support opposite-sign age slopes for overall accuracy, points won and blocks completed, with change points at 13.29, 13.66 and 14.55 years. Reaction time also passes a two-line test, but its turning point is 18.92 years; adults are faster than adolescents. Staying after a potential switch and asymptotic accuracy have significant quadratic age terms but do not pass the two-line U-shape test. Supplementary year-by-year plots resemble the age-bin plots, but are descriptive. Corrected pairwise tests do not establish that adolescents outperform every other group on every measure: for example, their overall-accuracy difference from ages 25–30 is not significant, whereas their points advantage is.
These qualifications retain the positive behavioral result while avoiding the stronger claim of one universal adolescent optimum. Accuracy, points and reversals are also partly dependent measures because reversal opportunities depend on accumulated rewards.
Models, estimation and validation
The winning RL model has four personal parameters: positive- and negative-outcome learning rates, inverse temperature and choice persistence. It starts both action values at 0.5, updates the chosen value toward observed reward and the unchosen value toward its complement, and uses the learning rate selected by the observed outcome for both updates. Factual and counterfactual rates are tied in the winning model. Persistence biases the next choice toward the preceding action, and inverse temperature controls the stochastic mapping from action values to choice.
The winning BI model also has four personal parameters: subjective reward probability, subjective switch probability, inverse temperature and persistence. It maintains a belief about which side is currently correct, updates that belief using feedback, propagates it through a switch model, then converts the belief into choice probability. The likelihood of reward on the wrong side is set to 0.0001 to avoid degeneracy rather than its experimental value of zero. The paper’s parameter-free BI comparison fixes reward probability to 0.75 and switch probability to 0.05.
Our qualification: a constant 0.05 hazard is not a literal representation of the full reward-count-dependent reversal schedule described in the Methods. Consequently, closeness to that value should not be treated as proof of an objectively optimal mental model for every trial of the actual task.
Parameters are fitted jointly across people using hierarchical Bayesian estimation in PyMC3, with two MCMC chains, 6,000 samples per chain and 1,000 discarded as burn-in. An age-free hierarchy estimates individual differences; a separate age-based hierarchy includes linear and quadratic age effects. Supplementary Table 6 specifies priors and hyperpriors, and Table 7 reports convergence statistics. Individual summaries are posterior means. The age-free estimates avoid directly building age into the displayed individual estimates; this does not make them assumption-free or remove shrinkage.
The winning RL model has WAIC 23,492 ± 201, versus BI 23,603 ± 200. Both reproduce reversal-aligned learning curves, outcome-conditioned staying and age-group patterns substantially better than their simpler versions. Simulate-and-recover analyses support parameter estimation, with hierarchical estimation improving recovery relative to independent maximum likelihood. Model-recovery simulations distinguish RL-generated from BI-generated behavior, although not perfectly. A concrete difference is that RL simulations become more likely to stay after two consecutive rewards than after one; BI simulations are already nearly certain after one diagnostic reward.
The evaluation is whole-task fitting, bias-corrected information-criterion comparison and model-based simulation/recovery. WAIC is used to estimate out-of-sample prediction error, but the paper does not conduct a chronological held-out-session forecast, held-out-person acquisition test, or own-history-versus-donor comparison. Its parameter-free BI baseline is not a strong fitted adaptive population control. Thus the results support useful individual descriptive models without isolating an incremental forecasting benefit of personal parameters beyond a shared model supplied with the same history.
Combining models and supplementary findings
Choice parameters correlate strongly between RL and BI. BI reward probability correlates with RL negative-outcome learning rate, while BI switch probability relates strongly to RL decision noise. Regressions using the other model’s parameters and pairwise interactions explain much of the fitted-parameter variance; positive-outcome learning rate is the notable exception. These are relationships among estimates from the same choices, not cross-task or future-choice predictions.
PCA of the eight standardized fitted parameters yields four components explaining 96.5% of their variance. Simulations at displaced component values motivate interpretations as behavioral quality, integration timescale, responsiveness to outcomes and positive-feedback updating. The authors combine age patterns in these components into an explanation involving adult-like performance quality, shorter integration timescales and distinctive reward processing in adolescence. This percentage refers to fitted-parameter variance, not variance in human behavior explained. The component labels are model-based interpretations, not independently measured biological faculties. Supplementary Table 16 even includes an illustrative RL positive learning rate of 1.19 for one displaced PC condition, outside the ordinary fitted Beta support; the component simulations should therefore remain qualitative probes.
Negative-outcome learning rate, BI reward probability and BI switch probability show visually nonmonotonic age patterns. However, none passes the supplementary two-line U-shape test. The age-based quadratic term is significant for switch probability; selected pairwise contrasts supply some of the evidence for the other parameters. Persistence and inverse temperature generally increase and then level off with age. Positive-outcome learning rate shows a different pattern, increasing in adulthood.
The supplement additionally covers prior reversal studies, sex-balanced age quantiles, recovery and MCMC diagnostics, additional outcome-history regressions, alternative age bins, corrected age-group comparisons, pubertal measures, alternative model variants, model discrimination, cross-model regressions and PCA simulations. Pubertal-development questionnaires and salivary testosterone broadly track age, but are highly confounded with it. Within-age analyses provide little robust independent evidence; some reported trends are uncorrected for multiplicity. Neither the main paper nor supplement establishes a hormonal or neural cause. Cross-sectional recruitment and unmeasured general abilities remain limitations, despite checks showing that the adolescent pattern persists without university participants and differs from other tasks completed in the same session.
Data boundaries and relevance to earlier personal experience
The paper links OSF 7wuh4 and SLCN model code. The separate code audit and data audit concern the currently released artifacts, not a demonstrated recreation of the paper’s exact analyzed rows. The code pin contains differing tutorial/main-length conventions and cleaning that deletes omissions before positional truncation. The public CSV lengths therefore must not be relabeled as authenticated raw main-task positions merely because the paper says 120 trials.
Our inference for Amadeus: this is a useful dataset and model family for testing whether a person’s earlier learning behavior helps forecast later choices beyond a history-adaptive population model. The authors explicitly distinguish context-dependent adaptation from universal parameter settings and favor the former interpretation. A useful personal forecast need not recover a fixed biological trait. Equally, a retrospective fit or cross-model parameter correlation does not answer the forecast question.
A prospective analysis could test earlier-task information at a frozen boundary, while all comparison models receive the same permitted target prefix. An own-versus-matched-other source comparison would ask whether correct personal linkage adds information; a shared-history-versus-specialized-update comparison would ask whether changing the update rule adds beyond carrying personal history through a common rule. If only earlier-task behavioral summaries are authenticated, a conditional forecast from those summaries into the cleaned B export is a legitimate, narrower question. It tests compressed earlier-person information, not transfer of a fully reconstructed trial-by-trial learning trajectory. Success would be evidence about this task sequence and observation process, with real-world continuation still open.