Li Ji-An, Marcus K. Benna and Marcelo G. Mattar, “Discovering cognitive strategies with tiny recurrent neural networks.” Nature 644, 993–1001 (2025), published online 2 July 2025. Final article; PMC copy; author code. The article is CC BY 4.0. This is our paraphrase and methodological assessment, not the authors’ text.
Read on 2026-09-29. Complete paper reading: final main article, Methods, declarations and references; Supplementary Information §§1.1–1.6 and 2.1–2.3, all Figures S1–S42 with captions and references; all three pages of the reporting summary. Main Figures 1–5 and supplementary figures were visually inspected, including plots whose equations or labels were imperfectly extracted. No standalone tables were present. Equations were read but not independently proved, and curves were not digitized. The publisher’s 22-page PDF includes the reporting summary; the separate SI PDF has a cover plus 44 numbered pages. The separately published 46-page peer-review correspondence was downloaded and text-extracted but remains unread; it is not treated as part of this full-paper reading. Cited papers were not recursively read. Local source and reading provenance.
What the paper establishes
This is positive evidence that flexible, low-dimensional recurrent models can predict held-out reward-learning choices better than several established cognitive models, while yielding interpretable descriptions of how model states change after actions and rewards. It is especially valuable because it separates the number of changing state variables from the number of fitted parameters. A recurrent model can have very few state variables while having substantially more parameters than a classical learning model. The comparisons chiefly match state dimensionality, not parameter count.
The study spans eight existing datasets and six task variants: reversal learning in monkeys and mice, two-stage learning in rats and mice, transition-reversal learning in mice, and three human tasks. The human samples comprise 1,010 people doing a 160-trial three-armed reversal task, 918 analyzed people doing a 150-trial four-armed drifting bandit, and 1,961 people doing a 200-trial two-stage task. The bandit source sample had 975 people; 57 were excluded for over 10% missing choices. The reporting summary also documents exclusion of two rats with fewer than 2,000 trials. No new participants were collected and no prospective sample-size calculation was performed.
Models predict the next action from preceding actions, states and rewards. Most use gated recurrent units, with fixed learned weights governing the changes in a recurrent hidden state. The Methods also specify switching GRUs, switching linear networks, and the cognitive comparators: model-free and model-based reinforcement learning, Bayesian inference, forgetting, perseveration and related variants. In the richer animal datasets, one to four recurrent variables often achieve the best or statistically indistinguishable performance. Human results require more qualification: two- to four-unit networks improve on corresponding cognitive models, but five to twenty units generally perform best, and large networks can be similar. Non-significant differences from a larger network are the paper’s operational dimensionality criterion, not a formal equivalence test or proof of the exact biological state dimension.
For animal fits, sessions are divided into approximately 150-trial blocks, with nested ten-fold evaluation: an outer test fold and inner training/validation folds. Networks use Adam, recurrent-weight L1 regularization, multiple initializations and validation-based early stopping. The reported test losses are trial-weighted. Direct individual fits need roughly 500–3,000 trials before recurrent models outperform the cognitive comparators. These are held-out blocks from the same animals, not a demonstrated chronological transfer to a later session or new animal.
Distillation and the human evaluation
The proposed remedy for limited individual data is a large population teacher followed by a small individual student. The teacher learns from multiple subjects; the student learns the teacher’s conditional action probabilities along one subject’s trajectories. The general architecture in Figure 2a and Methods includes a learned subject-ID embedding. In a representative mouse, distillation lets a small student beat the best equal-dimensional cognitive comparator with about 350 trials, whereas fitting a small network alone needs about 3,000. Figure 2b compares solo, teacher, student and cognitive-model performance. This supports the usefulness of population information for individual prediction and compression.
The main human results use interspersed within-block held-out labels. The random training/validation/test trial counts are 120/20/20 for reversal, 110/20/20 for the bandit, and 150/25/25 for two-stage learning. Each model receives the whole block as its sequential input; different losses score different trial indices. Thus predictions are based on recurrent histories, but fitted parameters can use training labels chronologically later than a test label. Actual test actions also become inputs to later trials. The target action is not simply fed into its own prediction. These distinctions matter: the procedure evaluates held-out choice prediction within an observed trajectory, rather than freezing an earlier personal model and forecasting a later session.
Supplementary Figure S40 tests this split on choices generated by a known cognitive model. RNN test loss does not beat the generating model’s loss in that check. The authors interpret this as ruling out data leakage. Our narrower interpretation is that the simulation is a useful implementation check under its tested generator; it does not establish prospective chronological separation or rule out every form of dependence between fitted parameters and held-out trial inputs.
Supplementary Figure S41 instead holds out people in six folds. The teacher trains and validates on five folds. For each held-out person, a student is fitted to teacher probabilities on action-augmented versions of that person’s entire block, validated against teacher probabilities on the original block, and then scored against the person’s actual choices on that block. This is meaningful generalization of population-learned predictions to new people. However, the student’s fitted weights summarize the teacher over an already observed whole trajectory; they are not parameters acquired from an earlier session and then tested on a later session.
Our central distinction: the human figures show student RNN advantages over the specified cognitive models. They do not isolate an incremental human forecasting gain from fitting an individual student or ID embedding over the same strong, history-adaptive pooled teacher. A shared recurrent policy already adapts its state to each person’s own experience. Individual student parameters and heterogeneous fitted dynamics therefore cannot, by themselves, distinguish a stable personal learning rule from a useful compression of that shared policy on different histories. This is a limit on the contrast answered, not evidence that personal learning is absent.
What the interpretation and supplementary analyses add
One-dimensional phase portraits describe changes in action preference as a function of current preference and trial outcome. They reveal state-dependent learning and perseveration, and reward-induced movement toward indifference after rare transitions. Higher-dimensional vector fields and regressions describe interactions between action values. Examples include unchosen-value updates and an unrewarded chosen value moving toward the unchosen value rather than toward zero. Human analyses suggest reward utilities, nonzero reference points and transition-dependent updates. These are concrete, testable candidate descriptions rather than merely high-level labels.
The supplement expands this with ground-truth simulations, data-efficiency curves, individual and fold comparisons, direct behavioral checks of state-dependent perseveration, switching-linear models, and detailed task-specific regression coefficients. RNN-inspired classical models improve the reported human test losses: 0.477 versus 0.602 in three-arm reversal, 0.796 versus 0.815 in four-arm bandits, and 0.448 versus 0.535 in two-stage learning. These are model-discovery results within the study’s evaluation framework; they are not an independent prospective replication of the newly proposed mechanisms.
Ground-truth simulations show recovery of the tested reinforcement-learning and Bayesian dynamics. A behavior-feature classifier trained on simulated model classes provides another check that empirical data contain the proposed signatures. Recovery for those simulations supports the method, while leaving open alternative untested mechanisms with similar observable predictions. The authors themselves discuss an alternative in SI §2.1: apparent state-dependent perseveration could partly reflect unobserved trial-to-trial preference variation that the recurrent model infers from later observations. Likewise, agreement across overlapping cross-validation fits is useful estimation stability, not long-term individual test–retest reliability.
The final application distills a larger meta-reinforcement-learning agent into a tiny recurrent model. The learned strategy approximates Bayesian inference while retaining history effects that distinguish it from an exact Bayesian solution. This demonstrates the interpretation method on an artificial agent; it does not add human personalization evidence. The supplement’s discussion of intrinsic dimensionality distinguishes continuous state manifolds from arbitrary real-number encodings and explicitly leaves a general conjecture open. The paper also acknowledges that limited data or difficult embeddings can underestimate or overestimate the behavioral dimension.
Targeted source-code and data-access audit
The code audit is partial and read-only, pinned to tinyRNN commit a8fc159a2b2d451be6756ae244b29b7d991aac1f. It covers all six human experiment scripts, relevant distillation and trial-mask functions, human loader setup, README and paths. No scripts were imported or executed, no saved teacher/student checkpoint was fetched, and no binding from these configurations to the published figure artifacts was established. Downloaded-code hashes and additional source/access hashes are retained.
There is a material qualification to the paper’s general architecture: all six inspected human configurations set include_embedding: False. The interspersed student scripts select 50-unit teachers; the cross-subject scripts select 100-unit teachers for Gillan and Suthaharan and a 20-unit teacher for Bahrami. Thus the Methods’ general 20-unit, ID-embedding description should not be assumed to describe every released human configuration. Several scripts contain inactive or commented execution branches, so source inspection alone cannot settle precisely which settings generated each published plot.
The core distillation helper replaces training and validation choice targets with teacher probabilities, leaving actual test targets intact. Trial selection changes masks, while keeping the whole input sequence. Cross-subject student scripts use augmented held-out blocks for training and the original block for validation/testing, consistent with the Methods. These checks support the split interpretation above without constituting a numerical reproduction.
Data availability was checked at the file-listing level, without downloading participant rows:
- Suthaharan repository: the public tree at commit
c46611d1bf86749c52eadb4dea464f766c6db220contains all eight MAT filenames expected by the tinyRNN loader underdata/hgfPRL. - Bahrami OSF: the public OSF API lists
DataAllSubjectsRewards.csvand a task-description PDF. The tinyRNN loader expects a locally renamed4ArmBandit_DataAllSubjectsRewards.csv. - Gillan OSF: public API listings expose both experiments, their
twostep_data_study1/2folders and self-report CSVs. The loader combines these studies and includes machine-specific paths that would need adapting.
The OSF browser route was inconsistent, but public API metadata succeeded. Data-content integrity, usable participant counts, repeated-session linkage and dataset-specific reuse terms were not audited here. The original data papers have not been read as part of this task; only their repository listings were inspected.
Implication for the Amadeus question
Retain this paper as evidence that learned recurrent updates and population-to-individual distillation are valuable tools for predicting learning behavior. The next discriminating human test should separately ask whether earlier personal experience helps beyond an adaptive population model with the same currently available history, and whether specializing the update rule adds beyond carrying that earlier experience through a shared model. Freeze acquisition before the future test interval, match personal initial-state/bias flexibility, and compare own acquisition with an appropriately matched other person’s acquisition. A repeated-session dataset is better suited to that question than relabeling the interspersed single-block evaluation as prospective personal learning. Success would establish useful personal predictive information in the tested task; further controls would still be required to identify an intrinsic personal mechanism or generalize to life events.
Peer-review file (read 2026-09-29, main session)
The separately published 46-page peer-review file (papers/carry_on/jian2025_tiny_rnns.supplement3.pdf), downloaded
but left unread by the earlier reading, is now read in full: all 2,676 extracted lines, three reviewer rounds and two
author responses, with every figure page (response Figures R1–R15 and R1–R2) inspected visually. The reports are
reviewers' views and the responses the authors' claims; neither is independent evidence.
- The human analyses were added in revision. Referee 2 called the first submission's distillation treatment "much too superficial", tested only on one animal. The three human datasets, the interspersed split and the cross-subject split were introduced in response. The authors state the interspersed split's design: the whole block is the model input and random trial indices are the targets (120/20/20, 110/20/20, 150/25/25), chosen so the three sets share one distribution. This confirms the reading above that the human evaluation is within-block, not a forecast of a later session.
- Generalization was the referees' main doubt. Referee 1 asked for a test of generalization to an untrained condition and suggested describing the RNNs as low-dimensional descriptions rather than discovered strategies. The response points to nested cross-validation, model recovery, a "behavior-feature identifier" and near-one logit correlations across outer folds; these are within-dataset checks. Referee 2 noted that long-range dependencies and meta-learning across blocks undermine fold-based cross-validation in many tasks.
- Effect sizes are small. In round two Referee 1 put the gain at about 1% in chosen-action likelihood per trial. The authors gave, for reversal learning, cross-entropy differences of −0.008 [−0.012, −0.004] and −0.015 [−0.021, −0.010] (monkeys, d = 1, 2) and −0.011 and −0.005 (mice), calling other tasks larger.
- Novelty was narrowed. After Referee 1 related the three signatures to adaptive learning rates (Pearce–Hall, Nassar), trial-to-trial learning-rate variability (Findling 2019) and side bias, the authors reframed them as established principles appearing in these tasks. The variability account of "preference-dependent perseveration" is the alternative noted from SI §2.1 above.
- Other points: "automatic" was removed from the title; "tiny" became 1–4 units; the cross-model "dimensionality" comparison was defended with a continuity argument (Hilbert-curve construction) that Referee 2 still found hand-wavy; the code was released only after Referee 2 required it; 50-unit students performed about as well as 20-unit ones.
For our question this strengthens the existing conclusion: the published evidence does not include a chronological, later-session personal forecast, and its incremental gain of an individual student over the shared teacher is not isolated.