Kurisutina

Humans create more novelty than ChatGPT when asked to retell a story

Scientific Reports 14: 875 (January 2024), DOI 10.1038/s41598-023-50229-7, PMC10776760, open access (CC BY 4.0). Indiana University Bloomington and the Chinese Academy of Sciences. Provenance: papers/carry_on/breithaupt2024_retelling_novelty.provenance.json.

What was read

All 467 lines of the text derived from the PMC XML: abstract, all sections, Table 1 (a worked example), figure captions, statements and all 57 references. The converter checks that the text keeps every non-whitespace character of the XML. Figures 1–4 are images; they hold most results, so figure-only numbers are missing. The OSF analyses (osf.io/jr2py) and the source study's supplement (ref. 29) were not read.

Question

How does retelling a story in a chain (Bartlett's serial reproduction) differ between people and ChatGPT? What changes and what stays?

Method

  • Stories. 116 short stories (120–160 words; happy, mildly happy, mildly sad or sad, without explicit emotion words), written by crowd workers in an earlier study (Breithaupt, Li & Kruschke 2022).
  • Human chains. 348 people each retold three stories after at least 40 seconds of reading, "in your own words", for the next person. Three retellings per chain.
  • ChatGPT chains. "ChatGPT3", given the same instructions. Each step came from a different OpenAI account.
  • Measures.
    • Word count and parts of speech (nltk).
    • WordNet synsets: survival from the parent version and the share of novel synsets.
    • Age of acquisition of the words.
    • Human ratings of happiness, sadness, imageability and "affected me", analysed with a Bayesian trend model.

Results

  • ChatGPT condenses once and then repeats itself.
    • The first retelling is much shorter than the humans' first.
    • After it there are few changes: synsets survive more and few new ones appear. The authors: "a broken disc … stuck on repeating the same core".
    • Its word-count decline over retellings is shallower, and it varies less than the humans'.
  • People reinvent.
    • About 55–60% of synsets are new at each retelling.
    • Perspective shifts between characters. In Table 1's example, the lonely man's story becomes the wife's justification for leaving, with an invented quarrel over politics.
  • Emotion is stable for both. Happiness, sadness, imageability and "affected me" barely change across retellings; all slopes and compressions are within ±0.1. The authors' puzzle: humans change most of the words but keep the emotional core, which they propose anchors the inventions.
  • Language differs.
    • Humans use more verbs, adverbs and pronouns, twice the negations (1.01% against 0.47%), and earlier-acquired words.
    • ChatGPT uses more nouns, adjectives and prepositions, and raises the vocabulary's age of acquisition above the original.
    • Age of acquisition predicts which words humans keep (earlier words survive) but not which words ChatGPT keeps.
  • Negations are more frequent in sad stories for both (humans 1.13% against 0.88%; ChatGPT 0.58% against 0.36%).

Limits

  • Not the same task for the two. Humans retold from memory after the text was removed. ChatGPT necessarily had the text in its prompt, so its "retelling" is summarisation (my inference; the paper says only that the instructions were identical). The ChatGPT chain measures repeated summarisation, not memory.
  • One ChatGPT version, unspecified beyond "ChatGPT3"; default settings; no prompt variations. Only three retellings. Crowd-worker stories.
  • Human retellings were rated on AMT and ChatGPT's on Prolific, so rater pools differ.
  • Inconsistencies found:
    • Raters for the human retellings are given as 537 in the overview and 492 in the affect section. The discussion's 1,068 total equals 537 + 531 (my computation).
    • The earlier ChatGPT transmission-chain study (Acerbi & Stubbersfield 2023) found human-like content biases. The authors reconcile this by method, not by data.

What it means for Kurisutina

  • Question 3: humans re-author their stories; an LLM freezes them.
    • Each human retelling keeps the emotional core and replaces most of the content, sometimes switching perspective. Serial reproduction is not recall of one's own experience, but it is the closest open measure of how a memory is re-told.
    • A replica that stores and re-tells its memories through summaries is likely to converge to a fixed version after the first pass. It would keep the gist, perhaps too faithfully, and stop re-interpreting (inferred).
    • Hirst found that people's memory of their own feelings at the time was the least consistent part. Here the emotional tone of a story is what stays. These are different constructs, which is another reason to score content, tone, one's own remembered feelings and perspective separately in the Q3 test.
  • A judge-free drift marker for Q2.
    • The model's voice shows as later-acquired vocabulary, more nouns and adjectives, and fewer negations and verbs.
    • A replica whose language moves from the person's baseline toward these markers is drifting toward the model's default. This can be measured without any LLM judge, which matters given the judge problems found in Abdulhai and Mooney.
    • Proposal: compute age of acquisition, part-of-speech ratios and negation rate on the person's own texts, and track the replica's distance from them over a run.
  • Retelling cycles in the replica's design. Reflection or consolidation loops (Park; MicroVerse) are retellings. If they only condense, the replica's memories will lose the person's re-interpretations, the kind of change that "carrying on" includes. Whether to let memories be re-authored (and how much) is a design choice to test against the person's own retelling behaviour.

Cross-references

  • summaries/carry_on/hirst2009_911_memory.md: human consistency and change over years.
  • summaries/carry_on/diamond2020_real_world_recall.md: forgetting with accurate retained detail.
  • summaries/carry_on/schacter2011_adaptive_distortion.md: constructive memory.
  • summaries/carry_on/park2023_generative_agents.md, summaries/carry_on/microverse2026_identity_drift.md: reflection loops.
  • summaries/carry_on/abdulhai2025_persona_consistency.md, summaries/carry_on/mooney2025_behavioral_coherence.md: judge problems that judge-free markers avoid.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.