Maria Katharina Eckstein, Sarah L. Master, Liyu Xia, Ronald E. Dahl, Linda Wilbrecht, and Anne G. E. Collins. The interpretation of computational model parameters depends on the context. eLife 11:e75474. Published article, PMC article and review correspondence.
Full reading completed/rechecked 29 September 2026. Context changes the interpretation and scale of fitted RL parameters, while some relative person differences generalize between tasks. This paper does not establish whether earlier personal history improves forward choice prediction beyond an adaptive population learner. A constant personal learning-rate trait is not a prerequisite for useful personal prediction.
Reading scope and retained evidence
Reread all 2,022 lines of the retained full-text extraction: abstract, introduction, all results, discussion, all Materials and Methods, appendices 1–8, references, ethics/data statements, editor's evaluation, decision letter, and full author response. Original XML/text are unchanged. Downloaded the publisher's 52-page PDF, including scientific appendices, and its text extraction. Visually inspected all 12 scientific figures (main Figures 1–5; Appendix 3 Figure 1; Appendix 4 Figure 1; Appendix 5 Figure 1; Appendix 8 Figures 1–4), all 14 scientific tables (main 1–6; Appendix 2 Table 1; Appendix 5 Tables 1–4; Appendix 6 Table 1; Appendix 8 Tables 1–2), and all four author-response images. All panels, axes, legends and table continuations were included. Tables were viewed on rendered PDF pages 7, 8, 9, 15, 16, 37, 43, 44, 45, 46, 51 and 52. This is full scientific text plus visual figure/table coverage, not a claim that each text-only PDF page was separately viewed as an image.
The separately listed supplement is the five-page MDAR reporting checklist, downloaded, read in full, and visually inspected on all five pages. The scientific supplements are the eight appendices within the article. MDAR states that the study was not preregistered or replicated and points to the exclusion, ethics, data and code sections; it adds no predictive experiment.
Retained sources: XML-derived text, publisher PDF, publisher text, MDAR PDF, visual artifacts, and provenance with exact coverage/hashes. The prior summary is preserved unchanged. Original task papers are identified below, but were not opened or read in this assignment.
Question, sample and chronology
The authors distinguish generalizability of parameter values or relative person differences across tasks from interpretability: whether a parameter consistently isolates a specific, distinct cognitive process. The same developmental sample completes three tasks during one laboratory visit. Analyses relate fitted whole-task parameters, behavioral summaries and age. Regression predictions are not forecasts of unseen later choices, and task-to-task directions are generally analyzed symmetrically rather than chronologically.
Methods report 312 tested, nine removed for age/missing-age reasons, and 303 remaining before task completion/performance exclusions. The complete-task analytic sample is 247: 143 ages 8–17, 51 undergraduates ages 18–28, and 53 community adults ages 25–30. Abstract and Study Design instead say 291. Task-specific completion counts also do not transparently reconcile with the final intersection: 45 community adults are said to complete C, but 53 appear in the final all-task sample. Raw membership must resolve these discrepancies. Performance exclusions disproportionately remove younger people and must not automatically define eligibility for a new forecast when they depend on future target behavior.
Published Methods, “Testing procedure” (PDF pp. 22–23), explicitly specifies an excluded four-choice reversal task first, then C → break → A → B. Ages 8–17 provided saliva and took a snack break after C. Order was identical for all participants; visits lasted 60–120 minutes. The author response repeats C→A→B when ruling out B-to-C strategy transfer. C is thus the earliest included task, not the person's first learning experience that day. Context, practice and fatigue are confounded with task position; chronological prediction is possible in principle, but their causal effects cannot be separated here.
| Task | Observations/actions | Context and duration | Winning model described here |
|---|---|---|---|
| A, butterfly | Four presented stimuli, two visible responses | Stable stochastic associations, 120 trials in four blocks, 10–20 min | Positive learning rate, inverse temperature, forgetting; negative rate fixed zero |
| B, stochastic reversal | Two boxes/actions, latent correct side | Correct side rewarded 75%, other side 0%; performance-dependent unsignaled reversals; 120 analyzed trials, 5–15 min | Positive/negative rates, temperature, persistence, counterfactual updates sharing corresponding rates |
| C, RL–working memory | One stimulus per trial, three responses, sets of 2–5 stimuli | Deterministic stable associations within block; ten blocks, 12–14 presentations per stimulus, 15–25 min | RL plus WM capacity, forgetting, noise and mixture; negative rate derived from positive rate and neglect bias |
A's probabilities are 70/30 in Methods/Figure 1, but 80/20 in Appendix 3. B's reversal range is 2–9 in Methods and 2–7 in Appendix 3. These are article inconsistencies, not extraction artifacts. Acquisition code from another revision cannot by itself authenticate a released file's settings. Tutorials, omissions, stimulus/action coding, block boundaries and analyzed trial ranges require raw-record and version checks. Figure 1 visually shows distinct observation/action structures and response deadlines, not interchangeable binary bandits.
Models and analysis, including every appendix
The shared RL form updates action values toward binary outcomes. A uses softmax and forgets absent stimuli toward 0.5. B adds previous-choice persistence and updates the unchosen action toward the complementary outcome. C combines RL values with a fast WM store, weights WM by capacity relative to set size, forgets its values toward 1/3, and adds uniform choice noise. C's negative rate is positive rate times a neglect factor, which also affects WM updating. Correspondingly named parameters can implement different mechanisms: A's forgetting acts on RL values, C's on WM; C's noise differs from temperature.
A/B use hierarchical Bayesian MCMC fits; C uses individual maximum likelihood because capacity is discrete. Appendix 4 describes six A candidates, a B sequence described as seven candidates, and six C candidates. Simulations based on fitted person parameters reproduce learning curves and recent-outcome patterns. These are checks of fitted behavior, not held-out-person or later-session prediction. Detailed fitting/recovery procedures are delegated to the original task papers.
Analyses include repeated-measures comparisons of absolute values; Spearman correlations; regressions using age and squared age on within-task standardized parameters; PCA on 15 parameters and 39 behavioral features; and ridge regression predicting parameters from parameters/behaviors. Ridge regularization and fold count (2–8 folds) are selected by R² with 100 repetitions, followed by refitting. The article does not describe an untouched outer evaluation of this complete selection procedure. Its reported cross-validation therefore cannot be substituted for our prospective, held-out-person incremental choice test.
- Appendix 1: definitions of accuracy, omissions, RT/variation, learning curves, stimulus-specific versus motor repetition, win-/lose-stay, delay effects, B's Bayesian-model and performance measures, and C's set-size/history measures. Most are whole-task summaries.
- Appendix 2: MDP/POMDP table, differing state/action counts and diagnosticity of feedback.
- Appendix 3: developmental task results: improving A performance, adolescent peak in B, reduced C set-size costs, and comparative behavioral age curves.
- Appendix 4: candidate models and fitted-model simulations for all tasks, including reproduced panels.
- Appendix 5: simulation ceiling construction, all four statistical tables and correlation intervals; the important correction to the old summary is below.
- Appendix 6: behavioral age regressions. Lose-stay trajectories differ between deterministic C and stochastic A/B; win-stay and response speed show greater consistency.
- Appendix 7: parameter trade-offs, unmeasured retest reliability, omitted choice-history mechanisms, task/model differences. This, not Appendix 8, contains the detailed limitations.
- Appendix 8: between-/within-task scatterplots, full feature correlation matrix, remaining PCA loadings/age curves, and two supplementary tables.
Findings and interpretation
Absolute fitted values differ substantially: mean positive rate is about .22 in A, .77 in B, and .07 in C; negative rates are about .62 in B and .03 in C (fixed zero in A). Noise expressed as inverse temperature is .095 in A and .33 in B; C's noise uses another parameterization. These differences accord with task demands but cannot be attributed uniquely to context rather than specification, estimation or order.
Relative person differences show partial generalization. Positive-rate and noise associations occur for A–B and A–C, with weaker/no associations for B–C. Age-controlled noise coefficients are .28, .19 and .039 respectively; positive-rate coefficients are .13, .23 and −.073. Negative-rate B–C association is −.12, p=.058 after age adjustment; the unadjusted scatterplot shows Spearman correlation about −.14, p=.03. These are different analyses. Forgetting is not significantly predicted A–C after age adjustment. Nonsignificance or similar age trajectories do not establish equivalence or reliable individual transfer.
PCA's first axis explains 25.1% and resembles broad performance/engagement; the next two explain 8.9% and 6.2% and contrast tasks. The figures confirm that parameters and performance/repetition measures mix on these axes. Remaining components have no sharp dimensional cutoff. Ridge relations cross parameter names: A's positive rate shares variance with C's WM capacity/mixture, among others. Figure 4's uneven R² and nearly unpredicted C WM parameters describe these fitted summaries, not a general inability to predict people across contexts. Behavioral relations to fitted parameters are more consistent than their scales, but shared variance does not uniquely identify a cognitive mechanism.
The simulation ceiling shares standardized differences, not identical raw parameters. Appendix 5 z-scores human parameters within task, averages each person's values across applicable tasks, and transforms the shared score back with each task's mean/SD. Simulated behavior is refitted/analyzed using the corresponding task-specific procedures. Shared ranking is compatible with large absolute task differences, visibly retained in Appendix 5 Figure 1A and Table 1. The main text's shorthand about identical parameters/no task differences must be qualified by this construction. Appendix 5 Table 2 also reports a forgetting task effect, p=.013, so broad prose about no differences for any parameter exceeds its table.
Human correlations fall below the simulated reference for most comparisons. Positive-rate A–B and noise A–C are not significantly below it; that does not prove equivalence to a true ceiling. The authors compare a human point correlation with the simulated correlation's bootstrapped 95% interval (1,000 BCa samples), rather than an interval for a paired difference carrying both uncertainties. The reference is conditional on fitted distributions, model families and construction. It diagnoses relative-parameter estimation; it neither rules out human model misspecification nor defines the attainable value of personal history.
Review exchange and limitations
Reviewers requested better fit/model-comparison evidence, reliability discussion, an FA sensitivity analysis, balanced treatment of successful generalization, a simulation ceiling, and MDP descriptions. The authors added these and changed the title. The four response images show small within-/between-task memory/rate associations, similar broad FA structure, and lose-stay histograms without clear bimodality. The last check does not establish universal instruction comprehension. Their objection to a giant common parameter model concerns identification/interpretation in these tasks; it does not rule out a shared predictive learner trained across people and conditioned on context/history.
All observations are from one visit in a developmental sample, not a factorial context experiment or same-task retest. Performance filtering changes the population. Parameter trade-offs and omitted history can generate or suppress correlations. A/C omit explicit outcome-independent persistence; B includes only one previous choice. Good aggregate simulation is not proof of unique mechanisms. Adaptation to task demands, modulatory processes, and parameters combining several processes are plausible accounts, not experimentally separated causes.
Visible reporting problems include a blank predictor in Table 5 and a forgetting label B–C in Appendix 8 Table 2 despite B lacking that parameter. Do not silently fill such entries or treat them as numerical ground truth. The article is CC BY 4.0 except explicitly credited material: Figure 5 is reprinted with Elsevier permission; Appendix 4 panels B/C have CC BY-NC-ND 4.0 credits. Retention for reading does not change those terms.
Implications for chronological personal prediction
Our question is whether C, or C plus A, improves later A/B choices beyond a strong population learner already adapting to the target's observed past. Confirmed linked trial records could support it; B cannot be used as earlier information when forecasting C. Keep context, stimulus/block identities, omissions and available actions explicit. Do not assume two- and three-action labels carry the same meaning.
First compare a shared learner with versus without permitted earlier personal history, measuring useful information. Then compare shared and specialized procedures given the same earlier and target history, measuring whether personal parameterization adds value. Personal adaptation through a shared model's state counts as useful personal prediction. A source-to-target mapping can be context dependent and learned from training people instead of imposing equal personal rates. Source identity controls test whether the person's own history is more useful than comparable other-person history, not intrinsic traits.
Population fitting and selection must exclude evaluated people. Individual adaptation must only use information available before each forecast. Whole-C features are available before A/B; whole-target summaries, fitted target parameters, future performance exclusions and future reversal labels are not. Order/practice may aid prediction without identifying a causal context effect. Observed next-choice prediction is also narrower than autonomous continuation or reactions to arguments/life events.
Availability is separate from reading. The article/MDAR link data and a notebook; that does not establish that linked trial-level records are public. Parallel repository/release auditors, not this reading, determine readiness. On 29 September they reported summary CSVs in the joint release and mismatches between current acquisition code and paper trial counts/probabilities. Do not substitute parameter correlations for forward choices or claim version correspondence without raw evidence.
Original-task references and public leads
These are metadata from this paper, not additional completed readings by this worker.
- A: Xia L, Master SL, Eckstein MK, Baribault B, Dahl RE, Wilbrecht L, Collins AGE (2021). Modeling changes in probabilistic reinforcement learning during adolescence. PLOS Computational Biology 17:e1008524. DOI; listed OSF dataset.
- B: Eckstein MK, Master SL, Dahl RE, Wilbrecht L, Collins AGE (2022). Reinforcement learning and Bayesian inference provide complementary models for the unique advantage of adolescents in stochastic reversal. Developmental Cognitive Neuroscience 55:101106. DOI; listed OSF dataset.
- C: Master SL, Eckstein MK, Gotlieb N, Dahl R, Wilbrecht L, Collins AGE (2020). Disentangling the systems contributing to changes in learning during adolescence. Developmental Cognitive Neuroscience 41:100732. DOI. This article's reference misspells “Distentangling”; its generated-data list provides no separate C dataset URL.
Joint data: OSF h4qr6. Analysis:
SLCN notebook.
The paper identifies Software Heritage revision 4fb5955c1142fcbd8ec80d7fccdf6b35dbfd1616.
Reading-history correction
The earlier statement that personal updating was answered “yes, but not as a portable trait” was too strong: this paper does not estimate added prospective choice-prediction value beyond population learning. Nor does it establish success for every similar context and failure for every dissimilar one. The prior 29 September targeted follow-up also misidentified the limitations appendix and treated a synthetic-recovery analogy as a prerequisite. That history remains in the archived prior summary; it is superseded, not a reason to resume the suspended Higashi work or add a new synthetic gate. This full reading supports a context-aware human predictive design, subject to raw-data availability, without a fixed-trait requirement.
Current records: next-context direction, completed human retest, and Cawley full reading.