Kurisutina

Associative intrusions and reported recollection

Research date: 2026-09-22.

Citation and status. Henry L. Roediger III and Kathleen B. McDermott. Creating False Memories: Remembering Words Not Presented in Lists. Peer-reviewed, Journal of Experimental Psychology: Learning, Memory, and Cognition 21(4), 803–814; July 1995. DOI; public author-laboratory PDF.

Acquisition and reading. Read all 12 pages: introduction, both experiments' full methods/results, general discussion, references, and the Appendix of word lists. The publisher-formatted scan has no useful first-page text layer; rendered all pages at 220 dpi and OCRed with Tesseract. Visually checked the title, all three figures, all three tables, the participant exclusion, and Appendix. No raw data or later replications were inspected; no separate supplement is referenced in the paper. OCR retains recognition errors and is a reading aid, not an authoritative transcription. Local PDF, OCR text, and provenance/hashes.

Source findings (approximately 150-word factual core). Experiment 1 tested 36 students with six selected associative word lists. Although instructed not to guess, they recalled the omitted critical associate on 40% of lists. Later recognition accepted critical lures on 84% of trials, versus 86% for sampled studied words and 2% for unrelated lures; 58% of critical lures received the highest confidence rating. Experiment 2 analyzed 30 students, counterbalancing lists across prior recall, arithmetic, and nonpresentation. Critical-lure recognition was .81 after recall, .72 after arithmetic, and .16 without the corresponding studied list. Respective proportions receiving a “remember” judgment were .58, .38, and .03. Instructions tied remembering to hearing the taped word, explicitly excluding merely remembering writing it during recall. Participants nevertheless reported recollection for absent items. These behavioral and subjective reports establish errors under this protocol; they do not locate an altered trace or isolate encoding, retrieval activation, source attribution, or a response criterion as the mechanism.

Methodological audit — additional source facts and their limits.

  • Designed materials and denominators. Experiment 1 deliberately selected strong intrusion-producing lists. Experiment 2 broadened materials but still constructed converging associates. The 40% and 55% false-recall figures are proportions of lists producing a particular lure, not percentages of all remembered words that were false. These are not population estimates for ordinary autobiographical recollection.
  • Test composition matters. Experiment 1's recognition test contained 12 studied and 30 unstudied words; each list-related block ended with its critical lure. Experiment 2 used 48 studied and 48 unstudied words in one randomly arranged order shared by all participants. Critical lures had a .16 baseline false-alarm rate when their lists were unstudied, versus .11 for ordinary words. Similar raw hit and lure rates therefore do not imply identical discriminability or evidence distributions.
  • Selection and stimulus repair. Experiment 2 excluded and replaced the participant who identified the lure-generating design; she was also the only participant with no critical false recalls. It scored 21 of 24 lists after removing two lists implicated in accidental presentation of critical lures and a third to balance sets. The Appendix explicitly replaces two words; it is not the exact original stimulus record. The exclusion limits claims about informed participants and warrants sensitivity analysis, without establishing that it explains the overall effect.
  • Different comparisons answer different questions. Counterbalanced recall-versus-arithmetic conditions compare activities following exposure. The paper does not report a separate inferential test for the .81-versus-.72 critical-lure recognition contrast; aggregate means do not supply its paired uncertainty. Table 3's later recognition conditional on whether a particular word was previously recalled is observational; the authors explicitly acknowledge selection effects. It cannot isolate the causal effect of producing that word. Nor does the experimental comparison isolate retrieval from all differences in interference or rehearsal between the two activities.
  • Confidence, reports, and mechanism. Experiment 1 measured graded confidence; Experiment 2 measured old/new followed by remember/know judgments for accepted items. The .58 “remember” figure is unconditional across tested critical lures; .58/.81, approximately 72%, is conditional on calling the lure old. Neither is a direct measurement of a uniquely identified recollection process. The paper discusses several model families but fits none; possible recognition-test priming and source confusion are explicitly acknowledged.

Interpretation for extraction — our reasoning. An interview answer can be authentic evidence about a person's present recollection while being incorrect evidence about the original event. Source records should distinguish an observed event, a person's later report, and an inference generated from related material. A repeated account does not become independent corroboration merely because it is now vivid or confidently endorsed. Moreover, the interview is part of the later history: a system forecasting future reports needs to know which earlier responses were elicited and potentially rehearsed.

The experiments do not demonstrate wholesale rewriting of an autobiographical trace. Association-driven generation, mistaken source attribution, and a decision policy can yield similar accepted-item counts while making different predictions under new cues or instructions. A person model needs prospective discrimination among these alternatives, not a declaration that it has reconstructed the hidden mechanism. The study also provides no estimate of how often vivid natural-life recollections are wrong.

Falsifiable acquisition and transfer test — proposal. Use controlled, neutral episodes with a recorded presentation history and plausible but absent details. After a common initial report, assign comparable item families to further free retrieval, re-exposure to that initial report, or an unrelated activity, counterbalancing family assignment. This tests the effect of an additional retrieval attempt; a separate group without the initial report tests its influence. Measure later event/source judgments, confidence, and actual choices that depend on the remembered details. Counterbalance recognition order, since the later test can itself supply cues. Preserve participant and item exclusions transparently; analyze all assigned participants, including those who detect the task structure.

Give recipient models identical acquired observations and comparable computation, varying whether they retain event/report/inference provenance. Keep unknown event truth for evaluation rather than inserting it into extraction inputs. Compare generic associative predictions with models receiving personal association observations, while matching exposure and interview history. Personalization succeeds only if it improves held-out predictions beyond generic associations and history, including source errors rather than merely average recall.

Score event accuracy, forecasts of the person's reports, and the recipient's own subsequent choices separately. A recipient that corrects the person's false belief may improve event accuracy while reducing behavioral fidelity; one that reproduces the report but makes unrelated choices has not transferred its functional consequence. The falsifiable engineering prediction is that preserving provenance helps retain both distinctions: what evidence supports the event and what the person currently takes to have happened. That result would support a representation useful for the tested functions, not prove a specific biological trace mechanism.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.