arXiv 2608.03044 v2 (31 August 2026; v1 4 August 2026, not read), cs.CL, CC BY 4.0. Queen's University, University
of Toronto. A short paper (5 pages of main text plus an appendix); no sign of peer review.
Provenance: papers/carry_on/griefalbert2026_emulate.provenance.json.
What was read
All 10 pages of v2: main text, limitations, ethics, references, and appendix A.1–A.5 (the LLM judge, all prompt scaffolds, metric definitions, per-condition Tables 5–6, first-token results). Figures were read from their captions and extracted labels.
Question
Why do studies of language models as simulated survey respondents disagree? The authors separate two tasks:
- emulation: the model answers as an individual respondent; many sampled answers make up a distribution;
- estimation: the model is asked directly for the distribution.
They then compare base and post-trained versions of the same models on both tasks.
Method
- Data. Pew American Trends Panel wave 54 (economic opinion, fielded September 2019): 59 four-option items. Refusals are dropped and the human distributions renormalised.
- Conditions (7). Unconditioned, plus Democrat, Republican, Very Liberal, Very Conservative, Upper-income, Protestant.
- Models: three matched base/post-trained pairs. Qwen3-14B-Base / Qwen3-14B; Olmo-3-1025-7B / Olmo-3-7B-Instruct; Olmo-3-1125-32B / Olmo-3.1-32B-Instruct. Claude Opus 4.6 is a reference for estimation only.
- Emulation: open response.
- The model answers without seeing the options, and an LLM judge maps each answer to an option. The judge agreed with one team-member annotator at κ = 0.66 on 200 pairs, and κ = 0.72 on the 146 where both assigned a letter.
- Base models get an INTERVIEWER/PARTICIPANT scaffold, with the demographic given as an earlier answer ("I am a Democrat.").
- Post-trained models get a system instruction to simulate a participant, sampled at temperature 1.5.
- 100 generations per question.
- Why not the first answer token's probabilities: the authors avoided first-token extraction because of positional bias (Wang et al. 2024; Tjuatja et al. 2024).
- Estimation. The model returns a JSON distribution over the options. Base models complete a partly written distribution; post-trained models get an instruction and are sampled greedily.
- Metrics. Total variation distance (TVD) and Wasserstein-1 over the ordered options.
Results
Average over the 7 conditions (Table 1; lower is better):
| Model | Estimation TVD / W | Emulation TVD / W |
|---|---|---|
| Qwen3-14B-Base | .236 / .492 | .263 / .457 |
| Qwen3-14B | .213 / .422 | .458 / .703 |
| Olmo-3-1025-7B | .278 / .605 | .248 / .415 |
| Olmo-3-7B-Instruct | .285 / .603 | .372 / .573 |
| Olmo-3-1125-32B | .223 / .460 | .274 / .443 |
| Olmo-3.1-32B-Instruct | .185 / .347 | .427 / .643 |
| Claude Opus 4.6 | .142 / .267 | — |
| uniform | .271 / .582 | — |
- Base models emulate better, in every model and condition. On Wasserstein, base emulation beats uniform in 20 of 21 model × condition cases; on TVD in only 11 of 21.
- Post-trained outputs collapse. They are 8–14× more self-similar at the bigram level (Table 4: bigram Jaccard 0.14–0.24 vs 0.014–0.020 for base). Higher temperature did not fix it.
- Group structure. Base models reproduce the human pattern of differences between the six groups (Spearman ρ = 0.61–0.75, all p < .02) but compress those differences to about 70% of human size. Post-trained models track them less well (ρ = 0.35–0.59; 0.35 not significant) and exaggerate them about 2× (mean ratio 2.01–2.24).
- Post-trained models estimate better. Estimation error falls with scale; the 7B pair is an exception. Base models also estimate at or better than uniform in most settings, so the distributional knowledge exists before post-training.
- First-token extraction (appendix A.5, Qwen3 0.6B–14B).
- Scaffolds were chosen to maximise probability on the answer tokens; the base prompt ends "My choice is Letter:".
- Base beats post-trained at every scale, best at 8B.
- The 14B base model falls to near-uniform, which the authors attribute to positional priors.
- The authors' account.
- Post-training steers the model into a single Assistant persona (they cite the Persona Selection Model and Lu et al. 2026).
- Emulating a group then becomes "second-order simulation": the Assistant plays the respondent, which squeezes diversity.
- The same Assistant produces good estimates.
- They name a simpler rival explanation they cannot rule out: post-training lowers output entropy with no persona structure involved.
Limits
- One survey, one domain (economic opinion, strongly partisan-sorted), and group-level distributions only. There are no individual respondents; the "persons" are demographic labels.
- The 2019 data are scored by models trained on 2024–25 text; the authors flag that the period is blended.
- Emulation depends on an LLM judge with κ 0.66–0.72, validated by a single team member.
- Post-trained emulation used temperature 1.5 and base models used their own scaffolds, so arms differ in more than training.
- Frontier base models are unavailable, so whether the tradeoff holds at scale is unknown.
What it means for Kurisutina
- The GSS pilot's post-trained arms use the weaker mode. The design has the post-trained models (Qwen3.5-9B, Qwen3.8-27B) answer in the first person and reads the first answer token: that is emulation, the mode where this paper finds post-trained models collapse. A fair test of them also needs an estimation read-out: ask for this person's probability over the answer codes, as the replay harness's Claude client already does. Without it, a base-model win in the pilot could be a mode effect, not a model effect.
- It supports the base model for the replica itself. A replica has to be the person, producing answers and text, not describe a distribution about them. By this paper, that is where base models are stronger, and the ~2× exaggeration of group differences is the "stereotyped persona" failure: reacting like a caricature of her group instead of like her.
- A risk for the base arm: positional bias in first-token read-outs. GSS codes are ordered numbers, not letters. They still carry a code-order prior that can dominate at some sizes (the 14B collapse). The pilot's empty condition exposes the model's default, and the population tables expose the truth. A code-order check (reversed codes on a subset) would isolate the bias directly.
- Individual prediction is a third task. The paper covers groups only; the pilot's per-person forecast sits between emulation (answer as her) and estimation (probabilities for her). Which mode wins there is exactly what the pilot can measure if it includes both read-outs.
Cross-references
- Lu et al. 2026 (
summaries/carry_on/lu2026_assistant_axis.md): the Assistant persona the authors invoke. - Peng et al. 2026 (
summaries/carry_on/peng2025_funhouse.md): post-trained twins regressing to a generic persona. - Namazova et al. 2025 (
summaries/carry_on/namazova2025_open_loop.md): a model fine-tuned on human choices predicted well but generated badly.