Mingyu Song, Yael Niv and Ming Bo Cai, Using Recurrent Neural Networks to Understand Human Reward Learning, Proceedings of the 43rd Annual Meeting of the Cognitive Science Society, pp. 1388–1394. Publisher record; author PDF.
Full reading completed 29 September 2026. Read all seven pages, including all main text, seven figures and their captions, the footnote, acknowledgment and references; visually inspected every page. Also read and visually inspected the linked one-page author poster. Retained paper, extracted text, poster and provenance. The eScholarship cover identifies CC BY 4.0. The author PDF has the same seven-page paper without the publisher cover/page furniture. No separate scientific supplement was located in the publisher record, the two authors' publication pages, or targeted title searches. The author page links a poster and conference video, not a methods supplement; the video was not viewed. No code execution or numerical reproduction was done, and cited papers were not recursively read.
What this paper directly establishes
The main comparison relevant to our research is a single trained population RNN versus an RNN with a learned participant embedding, both predicting human choices in a reward-learning task. The shared RNN already updates its hidden state using the participant's observed choices, stimuli and rewards. It is neither a fixed marginal choice distribution nor an average of separately fitted person models. Adding identity information improves held-out-game prediction modestly and reproduces individual behavioral distributions more faithfully. The strongest predictive improvement appears in the first two trials of a new game (p. 1392, Fig. 5).
This is evidence for useful person-specific information beyond the shared history-conditioned predictor. It does not isolate a person-specific learning-update rule from stable preferences, initial beliefs, exploration style or other effects of identity. That distinction follows from the architecture and timing of the gain; it is our interpretation, not a mechanistic result claimed by a controlled ablation in the paper.
Task, cognitive comparator and observations
The build-your-own-stimulus task has color, shape and pattern dimensions, each with three possible features. On each trial, participants select a feature or leave a dimension for the computer to randomize, giving 4³ = 64 possible choice configurations. An underlying game rule makes one, two or three dimensions reward-relevant, with one rewarding feature within each relevant dimension. Feedback is probabilistic and binary. In half the games participants know the number of relevant dimensions; in half they do not. Figure 1 shows the choice screen, response deadline below five seconds, 0.5-second stimulus presentation and 2.5-second feedback/intertrial interval (p. 1389).
There are 102 Mechanical Turk participants, each completing 18 games, three of each of the six condition types, with 30 trials per game. This implies 540 scheduled trials per participant and 55,080 scheduled trials overall; these totals are arithmetic, not independently audited retained-row counts. Performance improves within games, and more relevant dimensions make learning harder. Participants select more features in more complex known games; feature counts vary much less with complexity in unknown games (Fig. 2).
The strongest previously developed cognitive comparator is value-based serial hypothesis testing, combining feature-value learning with rule-hypothesis testing. It includes person-specific hypothesis-space and choice-policy parameters. Its geometric mean per-trial likelihood is reported near 0.22, versus 1/64 ≈ 0.0156 under uniform random choice. It captures important qualitative condition differences but misses quantitative behavior. The 2021 article delegates detailed task probabilities and cognitive-model specification to Song et al. (2020); it does not provide enough detail here to independently reconstruct that baseline's entire fitting and selection pipeline.
Shared and personalized models, splits and chronology
The shared model takes a game-start indicator, game type, previous choice, previous realized stimulus and previous reward. The last three inputs are zero on a game's first trial. Inputs are one-hot encoded and concatenated, followed by an LSTM, linear output and softmax with inverse temperature fixed at one. Training minimizes choice cross-entropy with Adam in PyTorch. The selected shared model uses 50 recurrent units, learning rate 0.001, batch size 10,000 and early stopping at epoch 28 (p. 1390, Fig. 3A).
Each participant contributes 16 games to training, one to validation and one to testing. Game type and game index are balanced across sets as far as possible. Hyperparameters are selected on validation performance, and the reported results use the test set. Training games are reported as augmented 1,024 times through feature/dimension shuffling. This is a game-held-out evaluation for already-known people, not a participant-held-out generalization experiment. The split does not establish that training games chronologically precede the test game: balancing game indices is not a forward-time split. Within a test game, the predictor conditions on actual prior choices/outcomes. Free-running simulations in later analyses answer a different question.
The personalized version maps one-hot participant identity to a learned embedding, concatenated with the other inputs and trained jointly with the shared network (Fig. 3B). Validation selects a three-dimensional embedding, otherwise the same reported learning rate, batch size and recurrent size, with early stopping at epoch 23. Thus the personal representation is a lookup learned from that person's other training games. It is not an encoder demonstrated to infer a new person's representation without fitting, and it is not a collection of independently trained RNNs. Because the two models are separately trained and selected, the comparison measures the performance of those two complete procedures; it does not hold their shared weights fixed.
The game-start input and zero previous-trial inputs are explicit. The text does not spell out every hidden/cell-state reset implementation detail or publish exact split membership. These remain reproduction questions, not assumptions to fill in silently.
Results and all figure evidence
Figure 2 compares human trajectories with cognitive-model and shared-RNN simulations. Both reproduce major reward-learning patterns, with the RNN closer on selected-feature trajectories. Figure 4 separates the predictive advantage: shared-RNN likelihood gains are primarily on switch trials, and it predicts stay versus switch better in both categories. This is not merely a bias toward switching. It also assigns more probability to choices with the correct number of selected features, narrowing the relevant part of the 64-choice space. Across people, improvement in this count prediction correlates with likelihood improvement (Fig. 4F: r = −0.57, p = 4.48×10⁻¹⁰). Within correct-feature-count switch choices, the reported true-choice probability ratio is 0.304 for the RNN versus 0.241 for the cognitive model. The remaining mechanism is left unresolved. The first trial is an explicit exception to the shared RNN's broad advantage over the cognitive baseline, which already has person parameters.
The paper reports an approximately 0.04 increase in geometric mean choice likelihood over the cognitive baseline, from about 0.22. Figure 5 shows a smaller additional gain from the participant embedding, concentrated in the first two trials. Later differences fluctuate around zero. The paper provides curves and participant standard errors, but no exact tabulated overall embedding gain or formal test of that overall difference. Approximate plot heights should not be promoted to precise effect estimates.
Figure 6 relates embedding principal components to three behavioral summaries: average selected-feature count (PC1, r = 0.88, p = 2.2×10⁻³³), dimensions changed on switch trials (PC2, r = −0.53, p = 1.3×10⁻⁸), and fraction of switch trials (PC3, r = −0.79, p = 4.6×10⁻²³). These are descriptive associations, not proof that an embedding axis identifies a cognitive mechanism. Free-running shared-RNN simulations reproduce means but underrepresent between-person dispersion. Embedding-conditioned simulations reproduce the distributions better. Feeding the real first two trials to the shared model recovers some variation, particularly feature count, but still misses other variation such as switch proportion. This tests what those two observations convey; it does not show that all longer target histories are insufficient.
Figure 7 addresses underfitting through simulation from the fitted cognitive model. As simulated training data grow from one to 1,296 times the original amount, RNN likelihood approaches the known generating model. Augmentation by dimension permutations (6), feature permutations (216), or both (1,296) improves effective data size by roughly tenfold, still short of the large-data result. The main-methods augmentation number is 1,024; the discussion's 1,296 is the full symmetry count. Their exact sampling relationship is not specified here. Augmentation assumes strategies are invariant to particular feature/dimension labels, an assumption relevant to any interpretation of personal preferences. The poster repeats the main results without additional training or split details. There are no tables or appended supplement in the seven-page paper.
Interpretation for person-specific learning research
The paper's proposed use of flexible RNNs is to diagnose predictable behavior missed by hand-built cognitive models. Its simulated-data study also shows that the fitted RNN is not a demonstrated information-theoretic ceiling. Better prediction provides an empirical reference, not a measured entropy bound. Embedding correlations may suggest hypotheses about individual strategies, but intervention or constrained model comparison would be needed to identify them.
For our question, preserve three separate claims. Population adaptation: a shared recurrent learner can use a person's target history. Personal prediction: training an identity representation on other games can improve prediction beyond that learner. Personal learning mechanism: the improvement comes specifically from how that individual updates after experience. This paper supplies direct evidence for the first two, with the third still unresolved. A first-trial improvement necessarily operates before feedback from the new game, so it cannot itself establish a different within-game learning update. Neither the embedding analysis nor the first-two-trial simulation supplies a static-bias-only or personal-initial-state-only control.
A useful direct human comparison would therefore retain a strong, trained, history-conditioned population model, add personal information using only permitted earlier observations, and separately examine gains before and after new feedback. A chronology-respecting held-out-game test and controls separating personal initial preferences from personal history effects would answer a stronger question than the paper's known-person, balanced-index split. These are research implications, not results obtained in this reading. We should not replace the shared adaptive baseline with a marginal frequency model or describe an average of personal networks as the trained population comparator.