Kurisutina

Reversal-learning reliability

Maria Waltmann, Florian Schlagenhauf, and Lorenz Deserno. Sufficient reliability of the behavioral and computational readouts of a probabilistic reversal learning task. Behavior Research Methods 54, 2993–3014. Published article, PMC record, public data/code.

Full reading completed 29 September 2026: all 22 main-article pages, including methods, discussion, limitations, references, all eleven figures and three tables; the complete supplementary DOCX, including all twelve tables and both figures. There is no separate main-paper appendix. Figures were visually inspected from the retained PDF and embedded supplement images, in addition to reading captions and text. The source manifest records URLs, hashes, coverage, and the limited code/data inspection. No fitting, simulation, or human-outcome analysis was performed for this reading.

The paper supports the existence of repeatable individual variation on this task, while showing that its estimated size depends strongly on the estimator and model. It does not test whether information from an individual's earlier session improves prediction of a later session beyond an adaptive population model. Its strongest estimates use both sessions during fitting. That distinction is central to our use of the evidence.

Forty healthy adults, aged 19–38, completed two counterbalanced versions of a probabilistic reversal learning task one week apart. Each session contained 160 binary choices. The two stimuli had complementary 80/20 reward probabilities, with five prescribed reversals after trials 55, 70, 90, 105, and 125. Feedback was sampled with replacement, so realized difficulty varied across people and sessions. Two participants were excluded for chance-level performance in one or both sessions and extreme accuracy relative to the sample; the analyses therefore use 38 people. The study was not preregistered.

The authors compared raw proportions/means with individual predictions from mixed-effects models, estimating the sessions separately or jointly. Behavioral outcomes were accuracy, overall/win/loss staying or switching, perseveration after consecutive losses, reaction times, and the difference between win and loss reaction times. Joint models nest session within person and use information from both sessions to estimate individual values. Retest statistics include absolute-agreement ICC(A,1), ICC(1), Pearson correlations, and model-derived variance measures. Odd/even-trial correlations quantify internal consistency; this split is not a forward prediction test.

For computational modeling, they fit 24 reinforcement-learning models: twelve variants using softmax inverse temperature and twelve using reinforcement sensitivity. Variants differ in chosen-only versus counterfactual updating, one versus separate win/loss learning rates, one versus separate win/loss choice or reinforcement sensitivities, and a weight on updating the unchosen option. Estimation is maximum likelihood, MAP with a Gaussian prior of mean zero and variance ten in fitting space, or empirical-Bayes EM-MAP with a multivariate population prior. They use repeated random starts and repeated complete fits, and compare models using integrated BIC. The main fits use analytic gradients and trust-region optimization; the supplement also examines quasi-Newton fitting.

The joint computational analysis concatenates the sessions but fits distinct session-specific parameters, resetting Q-values at session two. EM learns covariance among the parameters, including across sessions. The second session consequently informs both the group distribution and individual posterior estimates. This is a legitimate longitudinal estimation problem, but it differs from predicting the second session without seeing its choices.

The overall iBIC winner is full double updating with one learning rate and separate reinforcement sensitivities for wins and losses, DU-2ρα. Its reliability changes substantially with estimation method. For example, the learning-rate ICC(A,1) is about .16–.17 with maximum likelihood and .20 with MAP0. The table below gives the EM-MAP results from supplementary Tables 1 and 3; the last column is a model-derived Pearson correlation, not the same statistic as the two ICC columns.

Parameter in DU-2ρα Separate-session ICC(A,1) Joint-session ICC(A,1) Joint model-derived r
Learning rate .59 .83 .74
Win reinforcement sensitivity .64 .85 .86
Loss reinforcement sensitivity .42 .84 .86

Behavioral measures show the same broad pattern, with meaningful differences across measures. Accuracy's ICC(A,1) is .41 from raw proportions, .42 from separate models, and .66 from joint models; its model-derived ICC(1) is .53. Perseveration changes from .35 to .25 to .72 across those three approaches. Some stay/switch and reaction-time measures already have good reliability before joint estimation. The paper therefore does not support a blanket claim that all behavioral individuality is an artifact of pooling.

The strongest softmax-family model is a five-parameter weighted double-update model with separate win/loss learning rates and temperatures. It has uneven reliability: joint EM estimates include a negative win-learning-rate ICC(A,1), −.35, despite high reliability for the loss learning rate, .95. The supplement reports all model variants and confirms that additional parameters do not uniformly improve measurement. Analytic gradients do not change the main qualitative finding, although some more complex models are sensitive to optimization choices.

The recovery evidence addresses several distinct issues. Figures 4–5 simulate binary and continuous measurements with known across-session correlations: separate point estimates tend to underestimate reliability, joint point estimates can overestimate it, and model-derived variance estimates are closer to the generating correlations in these simulations. Figures 7–11 and main Tables 1–3 examine model fit, parameter/behavior recovery, and recovered reliability. Supplementary Figure 2 shows poor model recovery for the complex softmax winner, which can be better explained by the simpler reinforcement-sensitivity model. These are useful checks of estimation and identifiability within the selected model/task setup; they do not establish a personalized advantage on unseen later choices.

The authors explicitly acknowledge that pooling sessions can inflate correlations between point estimates. Their variance-based analysis is intended to account for estimation uncertainty, rather than to hide this issue. They also discuss the difficulty of obtaining reliable computational parameters from one new person, and propose normative samples for empirical priors. Small sample size, healthy young participants, one task version/family, and variable realized feedback limit generalization. They do not establish clinical prediction or neural validity; the discussion separately warns that model-derived fMRI regressors can create misleading associations.

For our question, a relevant new comparison would learn the population distribution and selection rules from other people, infer a held-out person's parameters using session one only, and score that person's session-two choices sequentially. A shared comparator should itself learn from the same session-two past choices and feedback; a non-adapting constant is inadequate. The incremental effect of first-session personalization should also be separated from a static preference or choice-noise adjustment if the claim concerns an individual learning rule. Neither session-two parameter fitting nor the paper's joint-session posterior can supply the personalized prediction arm without changing that prospective question. The paper's results motivate such a test but do not determine its answer.

The public OSF inventory is promising for that distinct question. The current RL_retest_OSF_v3.zip contains 80 raw MAT files whose names form 40 complete first/second-session pairs, MATLAB modeling and recovery scripts, R behavioral/reliability scripts, and the emfit toolbox. The 3,307,424-byte archive's SHA-256 matches the OSF API metadata. An ignored copy and inventory are retained under data/waltmann2022/. I read its README, Prepare_data.m, Main_script.m, and the winning model's one- and two-session likelihood functions, but did not audit the complete codebase. The preparation script identifies action, reward, accuracy, reaction-time, and session fields, and its exclusion rule can use either session. A forward test would need an explicit eligibility rule that does not select people on the future outcomes being evaluated, plus a separate audit of missing trials, session/version ordering, and record alignment. Filename pairing alone does not complete that audit. The main script also distinguishes its current ten-simulation recovery workflow from the retained legacy hundred-simulation workflow used for parts of the paper; this is a reproduction detail to preserve if the data are taken forward.

A subsequent authorized data-integrity audit loaded the MAT records only for schema, legal-code, task-metadata, and missingness checks. It confirmed 40 complete pairs, 12,800 scheduled rows, and 12,730 joint-valid choice/reward rows. The 70 omitted choices have matching missing rewards. One RT vector lacks its last element; no choice/reward trial was removed for this reason. The archive's expversion/type fields do not distinguish versions, while a filename suffix flips 1/2 across all pairs in a plausible counterbalanced-version pattern. Acquisition code certifying that suffix or explaining physical stimulus identities is absent. No predictive or personal consistency analysis was performed. The OSF project specifies CC BY 4.0; included toolboxes retain their separate licenses.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.