Kurisutina

Individual differences inside a shared learner

Read in full 29 September 2026. Excessive Flexibility? Recurrent Neural Networks Can Accommodate Individual Differences in Reinforcement Learning Through In-Context Adaptation, online 2025, volume 9:34–61. Publisher, retained PDF, provenance, supplement.

Coverage: all 28 main-PDF pages, Appendices A–D, equations, references and declarations; all eight supplement pages, including Text S1–S2. Visually inspected all twelve main and six supplementary figures and the key equations. No code execution or raw-data reproduction. Bibliography entries are not additional readings.

Findings and method

A pooled RNN can infer individual differences through recurrent state. Simulations compare common-parameter and individual-MAP reinforcement learners with shared networks; individual adaptation is possible with very small networks, but incomplete and sensitive to training. Stopping early or reducing capacity can also harm shared-process learning. An autonomous-simulation diagnostic probes whether inferred differences persist.

Human examples use Sugawara (143), Palminteri (20) and Waltmann (40). The first two split stimulus contexts across training/test sets within sessions; Waltmann trains on session one and scores session two. Individual models sometimes improve over common fits; superiority over the adaptive RNN is not established uniformly. The RNN receives previous choices/rewards, without participant-specific weights.

Crucially, Appendix A.5 selects human-data RNN checkpoints by test loss. Those comparisons are therefore not untouched evaluations. The reported score is exponentiated mean log likelihood, not action accuracy. The simulation diagnostic is heuristic and depends on the model fitted to generated trajectories. Matching an RNN does not establish that a cognitive model is complete.

Printed inconsistencies include Equation 8's choice-trace sign versus its interpretation, Equation 31's recurrent time index, and Supplement S5's learning-rate range versus its preceding definition. These need code reconciliation before reproduction.

Implication for our research

Our population comparator must be allowed to adapt to the observed person. Extra personal history can still improve a forecast, but its value must be measured against that comparator. We should separately test personal starting state/response policy and personal updating, using earlier acquisition and later human choices. A shared model receiving the same extended history also tests whether a special personal parameterization adds anything beyond that information. We should not suppress adaptation to create an easier baseline, or interpret a failed personal fit as absence of individual differences.

The research reset applies these distinctions. This paper motivates a direct empirical comparison; it does not justify another extended synthetic recovery study or establish an invariant personal learning trait.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.