Kurisutina

Not Yet AlphaFold for the Mind: Evaluating Centaur as a Synthetic Participant

arXiv 2508.07887 v1 (11 August 2025), cs.LG, CC BY 4.0. Osnabrück, Princeton, Brown. A short commentary (four pages) plus supplementary methods; no sign of peer review. Code and data: github.com/snamazova/centaur evaluation (as printed; not read). Provenance: papers/carry_on/namazova2025_open_loop.provenance.json.

What was read

All 13 pages of the extracted text: main text, data and code statement, contributions, supplementary information (prompt construction, evaluation, the three tasks with their models, data and procedures, the WCST example prompt) and the 15 references. Figure 1 and Figure S1 were viewed as rendered images, because the results are reported only there. Values below read from bar labels are exact; values read from curves are approximate (≈).

Question

Centaur (Binz et al. 2025, Nature) is Llama 3.1 fine-tuned on human choices from 160 experiments and proposed as a "participant simulator". Does it generate human-like behaviour, or only predict the next human choice?

The distinction the paper rests on

  • Predictive performance is forecasting each trial given the human's own previous choices and outcomes (closed loop, teacher-forced).
  • Generative performance is producing a whole sequence de novo: the model's own choices and their outcomes are appended to the prompt, so its context "evolves exclusively on the basis of its own behavior" (open loop).
  • The illustration: a repetition model that repeats its previous choice with a fixed probability predicts an adaptive person well, missing only the switch trials. Run on its own, it never reverses (following Palminteri et al. 2017).

Method

  • Models. Centaur-8B and Centaur-70B, base Llama-3.1-8B and -70B Instruct, and a domain-specific cognitive model per task. The text calls the Centaur sizes "7B"/"80B" in places and "Centaur-70B" elsewhere; the figure legend says 8B/70B.
  • Scoring.
    • Prediction: mean NLL of the participant's actual choice. For LLMs, the softmax is taken over the full vocabulary at the final token; for domain models, over the answer set. The authors note that following Binz et al.'s procedure disadvantages classic models, which are normally fitted per individual.
    • Generation: sampling at temperature 1, prompts and random generators reseeded per simulated participant, then compared on task-specific behavioural markers.
  • Tasks.
    1. Reversal learning. Two bandits at 80/20, reversed at trial 50 of 100. The task class is in Centaur's training data. The reference data are synthetic: a Rescorla–Wagner model (α = 0.5, β = 2.5, initial value 0.5), 32 seeds, not humans.
    2. Horizon task. Wilson et al. 2014, experiment 1: 4 forced choices then 1 or 6 free choices; in Centaur's training set. Human data from 31 participants; the RW model was fitted on 25 and tested on 6. Generation used all 31 participants' timelines.
    3. Wisconsin Card Sorting Test, computerised; not in Centaur's training set. 375 undergraduates (Steinke et al. 2020), with an 80/20 participant split. The domain model is a 4-parameter sequential-learning model (Bishara et al. 2010), fitted globally (r = 0.967, p = 0.656, d = 0.41, f = 0.05). Generation used the held-out participants' stimulus timelines.

Results (Figure 1)

Centaur-8B Centaur-70B Llama-8B Llama-70B Domain model Chance
Reversal, NLL (synthetic data) 0.4651 0.4629 1.4657 0.6460 RW 0.4253; repetition 0.5680 ≈0.69
Horizon, NLL (human) 0.4879 0.4835 0.6940 0.7410 RW 0.4162 ≈0.69
WCST, NLL (human, held out) 1.342 1.352 1.444 1.728 SL 0.871 ≈1.39
  • Reversal, generated. Before the switch all models chose bandit 1 about 80% of the time. After it:

    • Llama-70B reversed almost completely (≈0.05);
    • the RW reference fell to ≈0.2;
    • Centaur reversed only partly (≈0.3–0.4), with large variability across seeds.

    Figure S1 shows six selected Centaur-70B seeds. Some choose bandit 1 on every trial and never reverse. Others oscillate regularly between ≈0.4 and 0.6, which the authors read as random rather than reward-driven switching. So Llama-70B, which predicted worse, generated behaviour closer to the reference.

  • Horizon, generated. Humans' optimal-choice rate rises across the six free choices of horizon 6 (≈0.71 to ≈0.86), the move from exploring to exploiting. Centaur shows neither the horizon effect nor that rise; its curves stay ≈0.65–0.77. No model reaches the human curve.

  • WCST, generated. Human accuracy was ≈0.50, perseveration errors ≈0.07 and set-loss errors ≈0.02. The global SL model came close (≈0.49 / 0.10 / 0.05). Both Centaur models (≈0.28 accuracy, perseveration ≈0.22–0.29, set-loss ≈0.12–0.15) and both Llamas were far off. Centaur's predictive NLL on this unseen task (1.34) is barely below chance for four options (≈1.39).

  • Authors' conclusion. Centaur "achieves high predictive performance on tasks it was trained on" but fails to reproduce "the qualitative hallmarks of behavior that the tasks themselves were designed to measure". It is worse than a small cognitive model outside its training set. Prediction alone does not make a simulator.

Limits

  • A short preprint, not peer reviewed. Three tasks; figures only, no statistical tests; error bars are standard errors.
  • The reversal "human" reference is simulated by the same RW model that serves as the domain baseline, so that panel measures agreement with RW, not with people.
  • The horizon comparison uses 31 people and a 6-person test set. The RW predictive score is described as computed "between the simulated and actual free-choice trials", which is unusual and not fully specified.
  • Population-level comparison only: no individual is simulated as that individual, and generative fit is judged on group means of a few behavioural markers.
  • Scoring LLM NLL over the full vocabulary while domain models use the answer set may favour the domain models on prediction (the text does not discuss this).

What it means for Kurisutina

  • Next-step accuracy is not carrying on. A model can forecast Alice's next answer well when fed Alice's own history, and still behave unlike her once it runs on its own history. This is the replica's situation after hand-over. The replay harness's scored predictions are closed loop (the model sees the person's real evidence). Carrying on must also be scored open loop: the replica's own outputs and their consequences feed back, and the person's hallmark patterns are compared, not only per-item accuracy.

  • Open-loop failure takes the forms the drift question fears. Some runs lock into one response forever, others behave randomly, and the task's designed hallmarks (reversal, the horizon effect) disappear. For Amadeus the equivalents would be:

    • perseveration: never updating after a change in the person's world;
    • flattening toward generic responding;
    • losing the person-specific hallmark.

    Each can be named as a marker in advance and checked against the person over time.

  • Fine-tuning on humans in aggregate helped prediction, not generation. Centaur is a model trained on many people's behaviour, the "average participant" route. Its generative behaviour was less reversal-sensitive than the base Llama-70B's. Absorbing population behaviour can make a model a better forecaster and a worse simulator at the same time.

  • Outside the training distribution, a small explicit model beat the LLM on both counts. This supports keeping a simple person-level baseline (a calibrated, domain-specific model of the person) as a standing comparison. An LLM-based replica should have to beat it on open-loop fidelity, not just on prediction.

  • A rule to adopt. For each task or domain, pre-register the behavioural hallmarks that define the person there, and report closed-loop prediction and open-loop generation separately. Higashi 2026 shows the same pattern: strong teacher-forced prediction, and a weaker, coarser on-policy check.

Cross-references

  • Higashi 2026 (summaries/carry_on/higashi2026_individuality.md): teacher-forced evaluation, with on-policy checked only by block summaries.
  • Eckstein et al. 2022 (summaries/carry_on/eckstein2022_context.md): the source of the reversal task class in Centaur's training set (task B there).
  • Lu et al. 2026 (summaries/carry_on/lu2026_assistant_axis.md) and Li et al. 2024 (summaries/carry_on/li2024_persona_drift.md): drift when a model conditions on its own growing context.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.