PNAS 120(44): e2313790120 (October 2023), DOI 10.1073/pnas.2313790120, PMC10622889, open access (CC BY-NC-ND 4.0).
University of Trento and University of Winchester. Provenance: papers/carry_on/acerbi2023_llm_transmission.provenance.json.
What was read
All 301 lines of the text derived from the PMC XML: significance statement, abstract, all sections, methods and 27 references. The converter checks that the text keeps every non-whitespace character of the XML. Figures 1–2 are images, and their per-step proportions are not in the text. The SI appendix (full stories, coding schemes, interrater reliability, single-bias and single-story results) and the OSF data (osf.io/6v2ps) were not read.
Question
Human transmission chains preferentially keep some content (negative, social, threat, stereotype-consistent). Does an LLM that repeatedly summarises the same material keep the same content?
Method
- Model. ChatGPT-3 (version of 9 January 2023), default parameters.
- Prompt. "Please summarize this story making sure to make it shorter, if necessary you can omit some information: story". The output was fed back with the same prompt, three steps per chain, five chains per experiment. The study was preregistered.
- Material. Five published human transmission-chain studies:
- gender stereotypes (Kashima 2000);
- negative against positive (Bebbington et al. 2017);
- social against non-social (Mesoudi et al. 2006, the four stories merged into one);
- threat (Blaine & Boyer 2018);
- creation myths with several biases (Berl et al. 2021).
- Coding. Presence or absence of each original information unit, by the authors, with a blind third coder double-coding four of the five studies.
- Analysis. Linear mixed model of the proportion kept, by content type, with random effects for chain step and chain.
Results
The model shows the human biases (β is the difference in proportion kept):
- stereotype-consistent over inconsistent: β = 0.058 (P < 0.01);
- negative over positive: 0.117. Ambiguous details were resolved negatively (0.183), as in humans; e.g. a man "taking an old woman's bag" became a thief;
- social over non-social: 0.321;
- threat over negative and neutral: 0.523; negative over neutral: 0.070;
- negative, social and biologically counterintuitive content over other content in myths: 0.076.
Differences from humans:
- Gossip beat ordinary social information; in humans it did not.
- The stereotype-consistent advantage appears from the first step. In humans, early steps favoured inconsistent information.
As with people, most change happens in the first step; after that the model converges on a stable version.
Limits
- The task is summarisation with the text in the prompt ("make it shorter"), not recall. The human studies were memory-based retellings, so the parallel is in what gets selected, not in how memory fails. The authors note that a single summary is enough to produce the biases.
- One model version, one prompt (the authors say they did not test prompt wording), five chains, three steps. Chain step, with three levels, is a random effect, which the model cannot estimate well (my inference).
- The authors coded the outputs themselves; a blind coder double-coded four studies. Reliability figures are only in the SI (not read).
- Inconsistencies found: none of substance in the text read ("ChatGTP-3" typo; a truncated journal name in ref. 19). Breithaupt et al. (2024) found little human-like novelty in LLM chains; the two findings are compatible, because this paper measures retention of original units, not invention.
What it means for Kurisutina
- Question 3: a replica that consolidates by summarising will keep the population's priorities, not its person's.
- One summarising pass already favours negative, threatening, social (especially gossip) and stereotype-consistent content, and resolves ambiguity negatively.
- People show these biases on average. But a replica using the base model to condense its memories applies the average human's selection (and the training corpus's), not the person's.
- Over consolidation cycles its memories should drift toward that schema. That is Bartlett's schematisation, but with the model's schema (inferred).
- Where to look. The first summarisation is where most change happens (here and in Breithaupt). A Q3 audit should compare what the person keeps of a story or episode with what the replica's first consolidation keeps, per content category: negative, social, threat, stereotype-consistent, counterintuitive. The coding schemes of the five source studies are ready-made.
- Ambiguity resolution is a person-level trait worth measuring. Whether someone reads the man with the bag as thief or helper is plausibly individual. The model's default is negative. A replica of a charitable interpreter would need its slot to override this, and a test can check it (proposal).
Cross-references
summaries/carry_on/breithaupt2024_retelling_novelty.md: novelty and style in human against LLM retelling.summaries/carry_on/schacter2011_adaptive_distortion.md: schema-driven distortion.summaries/carry_on/diamond2020_real_world_recall.md,summaries/carry_on/hirst2009_911_memory.md: the human memory baseline.summaries/carry_on/park2023_generative_agents.md,summaries/carry_on/microverse2026_identity_drift.md: reflection and summarisation loops in LLM agents.