Kurisutina

LLMs struggle to simulate human belief updates in controlled environments

arXiv 2607.28347v1, 30 July 2026, CC BY-SA 4.0. Author repository, dataset, preregistration. Provenance and retained-source hashes: JSON.

Reading status, updated 2026-09-29

Main paper complete; complete scientific package still unavailable because the referenced Supplementary Information was not located. The 21-page main PDF was read in full, including all 875 extracted lines, Methods, references and back matter. Figures 1–3 and Tables 1–4 were visually inspected again; Figure 3's text extraction is corrupted, so its image was essential. The dataset card had been read in the earlier review. This update also read the pinned author README and baseline prompt template in full and visually inspected six additional author plot PDFs described below. No new human analysis or model execution was performed.

The bounded supplement search checked the original v1 PDF, HTML and complete seven-file arXiv source archive; current GitHub tree at 1f3016747beba458f4632e0929de8f105facf7ff; initial public-release tree at c5765ffcaf0e32d54e0a695087cb035115b05e47; release/branch listings; Hugging Face file listing; author publication-page metadata; and the public preregistration file listing. No explicit accessible SI link or complete SI document was found. The preregistration's file listing is empty, and its linked parent project returns HTTP 401. This establishes what these public locations provided, not that no supplement exists anywhere. No author contact or private access was attempted.

The standalone plots and prompt have identical Git blob hashes in the initial and current author trees. They are therefore original-release assets, but do not establish the contents or completeness of Supplementary Figures 4–7, their captions, or the temperature/OLMo methods and tests. The SI remains unread/unavailable; the preregistration's substantive protocol and the wider analysis code were not read. The author README now describes later analyses, including individual error metrics and a model-size sweep; those are not silently attributed to v1.

The pre-update summary and provenance are preserved byte for byte as pohl2026_belief_updates.pre_completion_2026-09-29.summary.md and .provenance.json under papers/carry_on/. Historical provenance also records two model-output files, data/pohl2026/simulated_initial.jsonl and data/pohl2026/ablations.jsonl, deleted locally on 2026-09-27. Their historical hashes remain recorded, but the files are absent and were not downloaded or verified again. The current hash check covers the new reading artifacts, snapshots and five still-present historical artifacts; it does not assert that all historical paths remain available.

Question and design

The paper tests whether a prompted LLM can reproduce a matched person's immediate stance after the same short persuasive exposure. It does not test whether observing that person's earlier belief updates helps predict a later update.

391 UK Prolific participants completed all three topics—universal basic income, penalty shootouts and weight-loss drugs—in random order. Each episode recorded initial stance on a five-category scale (−2 to +2) and familiarity; presented three curated Reddit comments; then recorded post-stance and a ranking of the comments' persuasiveness. Statement wording was randomly pro or con. Each topic had three comment packages: predominantly pro, predominantly con and more neutral. There were 1,173 person-topic episodes and 27 comments. Demographics and the ten-item Big Five questionnaire were collected after all three episodes. Table 4 describes the sample; the claimed representativeness concerns Prolific's age, sex and ethnicity dimensions, not every attribute or population.

The six main models were GPT-5.2, GPT-5-Mini, Claude-Opus-4.6, Gemini-3-Flash-Preview, Qwen3-32B and Llama-3.3-70B-Instruct. For each person-topic pair the model received the person's actual initial stance, familiarity, demographics/personality and the same comments. It produced a JSON post-stance, ranking and short explanation. The main temperature was 0.7 with provider defaults. This was prompted simulation with one released output per pair, rather than a trained personal model or repeated-session prediction. The prompt contains no earlier personal update episode.

The main comparison uses a two-sided paired sign-flip permutation test of the mean human–model post-stance difference, correcting across six models at α≈.0083. Additional analyses compare distributions and spread, average Kendall rank correlation, descriptive mixed models of post-minus-initial stance, persona ablations for three models, and simulations that generate initial stance too. Mixed models include participant intercepts and a topic variance component; |z|≥2 is explicitly a descriptive screen, not a test that two groups' coefficients differ.

Main results and visual evidence

Table 1 rejects the mean-difference null for Llama (p<.001), but not GPT-5.2 (.017), Claude (.058), Qwen (.082), Gemini (.52), or GPT-5-Mini (.75) after the six-model correction. This is a positive result about small aggregate mean discrepancies for several models. The paper calls these non-rejections faithful simulation; they do not establish equivalence or accurate forecasts for particular people.

Figure 2 and the distribution tests show that all six post-stance distributions differ from the human distribution (reported χ² p<.0001). Several models put too much mass on neutrality and too little at the extremes, with model-specific exceptions. Models generally change more frequently while their absolute-change distributions are narrower: the human SD is .87, greater than every model's (reported Brown–Forsythe p≤.002). Mean absolute change is not uniformly larger: Gemini's .42 is below the human .53. Figure 2's comment-level ranking spread means humans disagree with one another more about which comments persuade, whereas each LLM gives more consistent rankings across prompted personas.

Table 2's mean matched-person Kendall τ ranges from −.0361 to +.0196. This provides little average rank association in this design; an average near zero does not prove statistical independence or rule out subgroup structure. Explanations or rankings also do not authenticate the cognitive process causing an update.

Figure 3a's human associations passing the screen are initial stance, topic, familiarity and one employment contrast. Some LLM associations differ, including agreeableness, ethnicity, education or statement framing for particular models. A screened association in one group and its absence in another is not itself evidence that their effects differ. The bounded stance scale and subtraction of initial stance also mechanically constrain associations between starting position and change.

Table 3's demographic/personality removals have inconsistent effects across the three tested models and are assessed principally through mean discrepancies and their p-values. That result does not show that these attributes contain no predictive information, nor that another estimator could not use them. Figure 3b–m shows that replacing the real initial stance with a model-generated one yields distorted initial distributions; subsequent post-stances reject the mean-difference test for all six models (p<.0001). This supports grounding these particular simulations in observed initial stance. It is not an exhaustive comparison of possible personal memories or learning rules.

Figure 1 locates the experiment within a hypothetical repeated social simulation, but only the immediate update step is empirically tested. The discussion's “simulation drift” is a plausible risk from compounding errors, not an observed months-long or multiround human trajectory.

Additional original-release author assets read

The pinned prompt prompt_templates/templates2.txt instructs the model to simulate the named participant from the current topic, initial belief/familiarity, profile and shown comments. It requests a post-stance, ranking, explanation and general-public stance. It contains no history of previous personal updates. Its displayed personality-item descriptions omit big6, while a separate substituted personality block supplies values; this inspection alone does not establish which values the generating pipeline included.

Six standalone PDFs were visually inspected in full, with derived text retained even where font extraction is corrupted:

  • post_stance_temperature_ablation.pdf: nine panels for GPT-5-Mini, Qwen3-32B and Llama at temperatures 0, .7 and 2. Aggregate category shapes are broadly similar within each model across these settings. This does not establish temperature invariance of individual predictions or error rates. The main Methods reports more malformed/nonsensical high-temperature outputs for Qwen/Llama, but the missing SI prevents a complete audit of that evidence and exclusions.
  • post_stance_olmo.pdf: SFT, DPO and final stages for Olmo-3-32B-Think and Olmo-3.1-32B-Instruct. The Think distributions remain relatively close in shape across stages; the Instruct DPO/final panels visibly concentrate more mass away from extremes than SFT. There is no visually uniform improvement toward the human outline. These are aggregate plots, not a controlled estimate that post-training causes worse personal prediction.
  • delta_belief_by_topic_models.pdf and delta_by_initial_stance_and_topic_models.pdf: all six models' topic-specific update histograms and mean changes conditional on initial stance, compared with humans. The plots make topic-specific discrepancies and wide human variability visible; shaded SD bands are not uncertainty intervals for means.
  • delta_belief_by_topic.pdf and delta_by_initial_stance_and_topic.pdf: corresponding three-model persona-ablation panels. Removing profile fields produces different local shifts, without an obvious uniformly superior condition.

These assets partially illuminate the missing supplement's topics, but the summary does not assign them unverified SI figure numbers or claim a complete SI reading.

Prior local recomputation, retained rather than rerun

The earlier review used baselines.py on the retained main.jsonl. It reported reproducing Table 1 means to four decimals and permutation p-values within .003 in the normalized pro-topic frame. The following are that earlier computation, not results newly calculated or reverified during this completion audit:

Forecast Exact match MAE Predicted change Correlation with actual update
Persistence .654 .527 0 —
Population table, leave-one-out mode (mean for correlation) .647 .541 .061 .469
Gemini-3-Flash .517 .643 .416 .263
Claude-Opus-4.6 .465 .677 .556 .337
Qwen3-32B .416 .741 .616 .316
GPT-5.2 .400 .725 .645 .370
GPT-5-Mini .334 .819 .777 .315
Llama-3.3-70B .297 .903 .795 .310

Persistence retains the person's actual initial stance. The population table uses other participants with the same topic, comment package and initial stance; its mode and mean serve different reported metrics. On these reported scores, both baselines exceed every listed LLM in exact match and MAE; the conditional mean also has greater update correlation. This is a retrospective leave-one-out comparison within this dataset, not a prospective deployment estimate or a test of carrying personal update history. Human change frequency was reported as .346.

The earlier audit also found persistence's mean discrepancy about .118 with p<.001, despite its better individual scores, whereas GPT-5-Mini's mean discrepancy was close to zero despite much lower exact match. Mean error cancellation explains why an aggregate-bias test cannot stand in for personal forecast quality. The v1 main paper reports neither these baselines nor these individual error comparisons; the wider current repository has evolved since v1.

Limitations and discrepancies

The setting is one short exposure per topic in one sitting, with only three heterogeneous episodes per person. It does not test months of adaptation, life-event responses, updating about the same topic repeatedly, or an algorithm learning a transferable personal response rule. The main paper provides no equivalence test or individual-error benchmark; its repeated-person structure also requires care when interpreting episode-level uncertainty. A single stochastic output per pair cannot separate stable model differences from simulation noise.

Previously recorded discrepancies are retained with their evidential scope: the old released-data computation reproduced Table 1's signs as human minus LLM, opposite the caption's definition; this was not rerun here. The Table 1 caption misdescribes a p-value as the probability of a null hypothesis. Table 3's GPT-5-Mini narrative calls an increased signed discrepancy an improvement despite the table's stated criterion. Figure 2i's blanket claim of a higher average absolute update has the Gemini exception above. Claims of matched distributions in the abstract/Figure 3 caption need to be read against the explicit all-six χ² rejections.

Implications for the replica goal — our inference

The supported lesson is narrower than this summary's former claim that current state carries the forecast while traits and demographics do not. Observed initial stance helps these prompted simulators; their tested profile ablations give no consistent aggregate mean-bias benefit. This does not identify all information relevant to a person, establish a null effect of traits, or resolve whether earlier personal learning episodes help beyond an adaptive population model.

A direct personal-learning test would compare later-topic forecasts with and without causally earlier updates, while retaining the same target initial stance, familiarity, comments and available history for the adaptive population control. It would compare own history with appropriately matched other-person histories and score held-out people or episodes explicitly. Recovering recorded topic order is essential because topics were randomized. Only three topics make any cross-topic result narrow and potentially imprecise; an unsuccessful estimator would not establish absence of personal learning.

The prior baseline results justify demanding individual predictive gains over persistence and conditional population estimates. They do not justify imposing the observed 35% change rate on another setting, treating every update as erroneous, or declaring the broad replica objective impossible. Personal history, person-specific parameters, and stable biological traits remain distinct hypotheses here.

Related completed summaries

Jia/LISS, Lundberg, Kiley, Argyle, Santurkar, Chen, and Acerbi provide related context. No cited paper was newly opened for this update.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.