Kurisutina

The truth is out there: accuracy in recall of verifiable real-world events

Psychological Science 31(12) (2020), DOI 10.1177/0956797620954812, which is closed. Read as the accepted manuscript on PsyArXiv (ud63x; "Manuscript in press at Psychological Science; accepted June 5, 2020"; CC BY 4.0). Rotman Research Institute (Baycrest), University of Toronto, University of Pennsylvania. Provenance: papers/carry_on/diamond2020_real_world_recall.provenance.json.

What was read

All 1,474 lines of pdftotext -layout output (33 pages): abstract, all sections, Table 1, figure captions, survey text and references. Figures 1–5 are images; Figure 5's axis labels came through as garbled text. The supplementary material (detail corpora S1–S2, error types, internal-detail breakdowns, quantity survey Figure S5, Figures S1–S4) and the OSF audio guide were not read.

Question

Memory research stresses error and reconstruction. How accurate are freely recalled details of real, controlled experiences after days to years?

Method

  • Two staged real-world events with known ground truth.
    • Mask Fit Test: a scripted respirator-fitting procedure for hospital staff. 33 younger adults recalled it 29–1,243 days later (mean 267). Encoding was incidental: they were recruited weeks to years after it.
    • Baycrest Tour: an audio-guided art tour of a hospital floor with 27 target items and a scripted encounter with a confederate. Recall two days later: 22 younger adults (mean age 24) and 19 older adults (mean age 69).
  • Recall. The Autobiographical Interview, free recall plus general probes only, done once for the first time at test.
  • Scoring. Each clause is internal (episodic) or external. Internal details are then verifiable or not, and verifiable ones accurate or not.
    • Accuracy is output-bound: the share of verifiable recalled details that are correct.
    • Quantity is the share of a corpus of details recalled by at least two people: 61 for Mask Fit, 209 for the tour.
  • Survey. 54 memory scientists and 326 other academics estimated recall quantity and accuracy for a 30- or 70-year-old after 48 hours or two years.

Results

  • Quantity. People recall a small part of what happened: 23.75% of corpus details (Mask Fit), 21.18% (tour, younger) and 14.63% (tour, older). Quantity falls with age and with delay (r = −.43).
  • Episodic richness (internal/total details) is lower in older adults (d = 1.29) and falls with delay (r = −.42).
  • Accuracy is high: 93–95%. Mask Fit .95, younger tour .94, older tour .93. It does not differ by group (F(2,71) = .49), and it does not decline with delay up to about 3.4 years (r = −.19, n.s.; BF01 = 1.63).
  • Errors exist but are few. 56 of 74 people (75.7%) made at least one error, and 34 (46.0%) two or more. The types are sequence, labelling, perceptual, semantic and other.
  • Experts expect far worse. The median estimated accuracy for a young adult after 48 hours was 40% among memory scientists and 30% among other academics, with no difference between the groups.
  • Authors' account. In free recall, people filter out low-confidence content and coarsen the grain of their reports over time, trading detail for accuracy (Goldsmith, Koriat & Pansky). Forced-choice formats and suggestive procedures give much higher error rates.

Limits

  • Accuracy covers only verifiable details. Thoughts, feelings and transient features are unscored, and older adults produce fewer verifiable details (63% against 88%).
  • The events are novel and distinctive, low in emotion and social retelling. The authors say they are less prone to conjunction errors, wrong-event errors and contamination by retellings. The data were reused from earlier studies (no power analysis). The accuracy ceiling may hide age effects.
  • The quantity corpus counts only details recalled by at least two people, so quantity is an upper bound.
  • Inconsistencies found:
    • The survey recruited 68 memory scientists but reports 54. The exclusion rules given (all-"0–10%" answers, incomplete responses) are stated for the other academics, and the difference for scientists is not explained.
    • The final sentence is garbled ("they remain accurate as memory quantity and change").
    • Arithmetic checked: 56/74 = 75.68% and 34/74 = 45.95%; F(1,378) matches 54 + 326 respondents.

What it means for Kurisutina

  • Question 3: the human target is "forget most, keep what you keep accurately".
    • For ordinary, low-emotion experiences, people freely recall about a fifth of the details, and about 94% of what they volunteer is right, even years later.
    • Hirst's 9/11 results (consistency .57–.63, confident repetition of errors) apply to emotional, much-retold events.
    • Together: the person's memory differs by kind of event. Details of ordinary episodes are thinned but accurate. Emotional, rehearsed ones drift, with confidence.
    • The prior drawn earlier from Schacter and Nichols ("memory distorts") should not be applied to everything (inferred).
  • Predicted replica failure: too much, not too wrong (guess, to be tested).
    • A perfect-retention replica (EM-LLM style) will recall far more than the person.
    • A generative replica asked "tell me everything you remember" may pad with plausible details, so its output-bound accuracy falls below the person's 94%.
    • Both differ from the person. Scoring must separate quantity, accuracy and grain, as the Autobiographical Interview does. Accuracy alone would miss the first failure; quantity alone the second.
  • A ready protocol for the Q3 test.
    • A staged, verifiable event that the person experiences and the replica receives as a record, such as an audio-guided tour; the materials are public on OSF.
    • Free recall at 2 days and later; Autobiographical Interview scoring; a detail corpus for quantity.
    • It combines with St. Jacques & Schacter's reactivation phase on a similar museum tour. The grain shift is a distinct, scoreable hallmark of human remembering over time.
  • Monitoring is part of human accuracy. People withhold what they are unsure of. A replica needs an equivalent, either calibrated confidence or a threshold for saying "I don't remember", or it will not match the person's output even with the same underlying memory (inferred).

Cross-references

  • summaries/carry_on/hirst2009_911_memory.md: consistency and confidence for an emotional public event.
  • summaries/carry_on/stjacques2013_reactivation.md: museum-tour reactivation and false alarms.
  • summaries/carry_on/schacter2011_adaptive_distortion.md, summaries/carry_on/nichols2019_false_memory_tasks.md: distortion as a by-product; task-specific susceptibility.
  • summaries/carry_on/fountas2024_em_llm.md: perfect retention.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.