Kurisutina

Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation

arXiv 2608.03044 v2 (31 August 2026; v1 4 August 2026, not read), cs.CL, CC BY 4.0. Queen's University, University of Toronto. A short paper (5 pages of main text plus an appendix); no sign of peer review. Provenance: papers/carry_on/griefalbert2026_emulate.provenance.json.

What was read

All 10 pages of v2: main text, limitations, ethics, references, and appendix A.1–A.5 (the LLM judge, all prompt scaffolds, metric definitions, per-condition Tables 5–6, first-token results). Figures were read from their captions and extracted labels.

Question

Why do studies of language models as simulated survey respondents disagree? The authors separate two tasks:

  • emulation: the model answers as an individual respondent; many sampled answers make up a distribution;
  • estimation: the model is asked directly for the distribution.

They then compare base and post-trained versions of the same models on both tasks.

Method

  • Data. Pew American Trends Panel wave 54 (economic opinion, fielded September 2019): 59 four-option items. Refusals are dropped and the human distributions renormalised.
  • Conditions (7). Unconditioned, plus Democrat, Republican, Very Liberal, Very Conservative, Upper-income, Protestant.
  • Models: three matched base/post-trained pairs. Qwen3-14B-Base / Qwen3-14B; Olmo-3-1025-7B / Olmo-3-7B-Instruct; Olmo-3-1125-32B / Olmo-3.1-32B-Instruct. Claude Opus 4.6 is a reference for estimation only.
  • Emulation: open response.
    • The model answers without seeing the options, and an LLM judge maps each answer to an option. The judge agreed with one team-member annotator at κ = 0.66 on 200 pairs, and κ = 0.72 on the 146 where both assigned a letter.
    • Base models get an INTERVIEWER/PARTICIPANT scaffold, with the demographic given as an earlier answer ("I am a Democrat.").
    • Post-trained models get a system instruction to simulate a participant, sampled at temperature 1.5.
    • 100 generations per question.
    • Why not the first answer token's probabilities: the authors avoided first-token extraction because of positional bias (Wang et al. 2024; Tjuatja et al. 2024).
  • Estimation. The model returns a JSON distribution over the options. Base models complete a partly written distribution; post-trained models get an instruction and are sampled greedily.
  • Metrics. Total variation distance (TVD) and Wasserstein-1 over the ordered options.

Results

Average over the 7 conditions (Table 1; lower is better):

Model Estimation TVD / W Emulation TVD / W
Qwen3-14B-Base .236 / .492 .263 / .457
Qwen3-14B .213 / .422 .458 / .703
Olmo-3-1025-7B .278 / .605 .248 / .415
Olmo-3-7B-Instruct .285 / .603 .372 / .573
Olmo-3-1125-32B .223 / .460 .274 / .443
Olmo-3.1-32B-Instruct .185 / .347 .427 / .643
Claude Opus 4.6 .142 / .267 —
uniform .271 / .582 —
  • Base models emulate better, in every model and condition. On Wasserstein, base emulation beats uniform in 20 of 21 model × condition cases; on TVD in only 11 of 21.
  • Post-trained outputs collapse. They are 8–14× more self-similar at the bigram level (Table 4: bigram Jaccard 0.14–0.24 vs 0.014–0.020 for base). Higher temperature did not fix it.
  • Group structure. Base models reproduce the human pattern of differences between the six groups (Spearman ρ = 0.61–0.75, all p < .02) but compress those differences to about 70% of human size. Post-trained models track them less well (ρ = 0.35–0.59; 0.35 not significant) and exaggerate them about 2× (mean ratio 2.01–2.24).
  • Post-trained models estimate better. Estimation error falls with scale; the 7B pair is an exception. Base models also estimate at or better than uniform in most settings, so the distributional knowledge exists before post-training.
  • First-token extraction (appendix A.5, Qwen3 0.6B–14B).
    • Scaffolds were chosen to maximise probability on the answer tokens; the base prompt ends "My choice is Letter:".
    • Base beats post-trained at every scale, best at 8B.
    • The 14B base model falls to near-uniform, which the authors attribute to positional priors.
  • The authors' account.
    • Post-training steers the model into a single Assistant persona (they cite the Persona Selection Model and Lu et al. 2026).
    • Emulating a group then becomes "second-order simulation": the Assistant plays the respondent, which squeezes diversity.
    • The same Assistant produces good estimates.
    • They name a simpler rival explanation they cannot rule out: post-training lowers output entropy with no persona structure involved.

Limits

  • One survey, one domain (economic opinion, strongly partisan-sorted), and group-level distributions only. There are no individual respondents; the "persons" are demographic labels.
  • The 2019 data are scored by models trained on 2024–25 text; the authors flag that the period is blended.
  • Emulation depends on an LLM judge with κ 0.66–0.72, validated by a single team member.
  • Post-trained emulation used temperature 1.5 and base models used their own scaffolds, so arms differ in more than training.
  • Frontier base models are unavailable, so whether the tradeoff holds at scale is unknown.

What it means for Kurisutina

  • The GSS pilot's post-trained arms use the weaker mode. The design has the post-trained models (Qwen3.5-9B, Qwen3.8-27B) answer in the first person and reads the first answer token: that is emulation, the mode where this paper finds post-trained models collapse. A fair test of them also needs an estimation read-out: ask for this person's probability over the answer codes, as the replay harness's Claude client already does. Without it, a base-model win in the pilot could be a mode effect, not a model effect.
  • It supports the base model for the replica itself. A replica has to be the person, producing answers and text, not describe a distribution about them. By this paper, that is where base models are stronger, and the ~2× exaggeration of group differences is the "stereotyped persona" failure: reacting like a caricature of her group instead of like her.
  • A risk for the base arm: positional bias in first-token read-outs. GSS codes are ordered numbers, not letters. They still carry a code-order prior that can dominate at some sizes (the 14B collapse). The pilot's empty condition exposes the model's default, and the population tables expose the truth. A code-order check (reversed codes on a subset) would isolate the bias directly.
  • Individual prediction is a third task. The paper covers groups only; the pilot's per-person forecast sits between emulation (answer as her) and estimation (probabilities for her). Which mode wins there is exactly what the pilot can measure if it includes both read-outs.

Cross-references

  • Lu et al. 2026 (summaries/carry_on/lu2026_assistant_axis.md): the Assistant persona the authors invoke.
  • Peng et al. 2026 (summaries/carry_on/peng2025_funhouse.md): post-trained twins regressing to a generic persona.
  • Namazova et al. 2025 (summaries/carry_on/namazova2025_open_loop.md): a model fine-tuned on human choices predicted well but generated badly.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.