Journal of Experimental Psychology: General 138(2): 161–176, May 2009, DOI 10.1037/a0015527. Read as the NIH author
manuscript (PMC2925254, NIHMS219091). The file carries APA's notice that it is the final accepted manuscript, not
copyedited. Empirical, peer-reviewed; a multi-site study (first author at the New School for Social Research, New York).
Provenance: papers/carry_on/hirst2009_911_memory.provenance.json.
Author list: the sixteen names match the requested citation. The order of two differs: the manuscript lists Lyle before Lustig, while PubMed (PMID 19397377) and the requested citation list Lustig before Lyle.
What was read
- Read: all 885 lines of the text derived from the PMC XML. That covers the abstract, introduction, method, results,
discussion, the three footnotes, the acknowledgements, APA's manuscript notice, all 68 references, the captions of
Figures 1–3, and Tables 1–8 with their notes.
- In the text file, table footnote markers are joined to the numbers: ".572" is .57 with footnote 2 (p < .01). This reading was checked against the XML.
- Not read:
- Figures 1–3 are images; their captions were read.
- Figure 1 holds the mediation coefficients.
- Figure 2 holds the proportions of responses that stayed consistent, or were corrected, repeated or changed to something else. The text gives only two of them (82% and 44%).
- Figure 3 holds the media-coverage and accuracy z-scores.
- The surveys (about 17 pages each) and the coding manuals. The paper places them at http://911memory.nyu.edu, and that host did not resolve on 2026-09-25. The wording of the Survey 2 and 3 questions is therefore known only from the text.
- There is no supplement (PMC: has-supplement no).
- Figures 1–3 are images; their captions were read.
Question
- After the first year, does forgetting slow down or speed up? This is asked for two kinds of memory:
- flashbulb memory: the circumstances in which one learned of 9/11 (where, from whom, what one was doing, how one felt);
- event memory: facts about the attack itself.
- Are emotional reactions remembered as well as the other features?
- Do the same factors predict the retention of both kinds, and does their content change in the same way?
- What role do a community's "memory practices" play, meaning media coverage and conversation?
- Background, as cited in the paper (not measured here):
- Schmolck, Buffalo & Squire (2000) studied memory of the O.J. Simpson verdict. At 15 months, "a little less than 40%" of memories had no distortions and "only about 10%" had major ones. At 32 months, "only about 20%" had none and "over 40%" had major distortions. So forgetting accelerated.
- Diary studies show fast forgetting in the first year, then slower:
- Linton: "a much slower rate of forgetting of 6% for the next five years";
- Wagenaar: "a substantial decline of 20% in the first year for critical details and then a slower decline of approximately 10% for the next four years".
- Talarico & Rubin (2003), over up to eight months: flashbulb memories are forgotten like ordinary memories, but confidence in them stays high.
Method
-
Design. The same people answered surveys at three points:
- one week after the attack (recruited 17–21 September 2001);
- 11 months after (5–26 August 2002);
- 35 months after (9–20 August 2004).
The 11- and 35-month intervals were chosen to match Schmolck et al. while avoiding anniversary commemorations.
-
Recruitment.
- Places: Boston and Cambridge, New Haven, New York, Washington DC, St. Louis, Palo Alto and Santa Cruz.
- Routes: tables on campuses or in nearby neighbourhoods, and friends and acquaintances of lab members. The return rates are called comparable to surveys "without a monetary incentive or a follow-up query".
- New participants were added at Surveys 2 and 3, to check for effects of taking part earlier.
- Surveys 2 and 3 had paper and web versions. Answers did not differ between them (all p > .4), so the data were merged.
-
Sample.
- Table 1 lists 1,060 people with more than one survey and 2,186 with one (3,246 in all; our sum). The abstract says "more than 3,000".
- "38% of the respondents to Survey 1 completed Survey 2", and 18% completed all three.
- Most analyses use the 391 people who completed all three surveys.
-
Probes (Survey 1, Table 2).
- Flashbulb memory, six features:
- how one first learned of it (source);
- where one was;
- what one was doing;
- how one felt on becoming aware;
- the first person one communicated with;
- what one was doing immediately before.
- Event memory:
- the number of planes;
- the airlines;
- the crash sites;
- where President Bush was;
- the order of six events.
- Predictors:
- personal loss (Q12) and inconvenience (Q13);
- intensity of sadness, anger, fear, confusion, frustration and shock on a 1–5 scale. Survey 1 asked about "CURRENT FEELINGS": "At this moment, how strongly or intensely do you feel …";
- media attention and conversation (each rated 1–5);
- the percentage of waking hours spent on listed activities;
- demographics, including residence.
- Surveys 2 and 3 had two versions, in equal numbers.
- Half the participants rated their confidence in each flashbulb answer (1–5).
- The other half forecast how accurately they would remember it two years ahead (Survey 2) or seven years ahead (Survey 3).
- Flashbulb memory, six features:
-
Coding.
- A manual was built from 50 surveys. When 50 similar answers had been coded "other", a new category was added and the question recoded; that happened for 14% of the questions.
- 10% of surveys were coded twice. All kappas or alphas were above .80.
- Flashbulb consistency. Each feature scores 1 if coded the same as in Survey 1, else 0. The mean over the six
features runs from 0 to 1.
- S12 compares Survey 2 with Survey 1; S13 compares Survey 3 with Survey 1.
- The score is dichotomous, unlike Neisser and Harsch's graded 0–2 scheme.
- On 50 participants, the two schemes correlate r = .29 (S12, p < .05) and r = .38 (S13, p < .01). The forgetting analyses redone with Neisser–Harsch scores gave the same pattern.
- Event accuracy, against news accounts:
- planes and Bush's location score 1 or 0;
- airlines: .5 per correct carrier, −.25 per wrong one (penalty at most .5), with negatives set to 0;
- crash sites: .33 per correct site, −.16 per wrong one;
- order: the Spearman correlation with the true order, with negatives set to 0;
- the total is the mean of the five probes.
- The paper says "accuracy" for event memory and "consistency" for flashbulb memory. There is no record of the personal circumstances to score them against.
- Emotion memory is measured in two ways:
- the consistency of the open-ended feeling (Q4);
- per person, the correlation of the six emotion ratings between surveys, after z-scoring within each survey. This is meant to capture memory of "the relation among different emotions".
-
Statistics.
- The significance level was set at .01, because of the large sample and many comparisons.
- Cohen's d is reported with the benchmarks .20 small, .50 medium and .80 large.
Results
Flashbulb memory (Table 4; the 391 three-survey participants; ** p < .01 and * p < .05 against 11 months):
| Measure | Survey 1 to 2 (11 months) | Survey 1 to 3 (35 months) |
|---|---|---|
| Overall consistency | .63 (SD .20) | .57 (SD .23)** |
| Emotional consistency (open-ended feeling) | .42 (SD .49) | .37 (SD .48) |
| Correlation of emotion z-scores | .48 (SD .40) | .42 (SD .39)* |
| Overall confidence (1–5) | 4.41 (SD .64) | 4.25 (SD .93) |
- Forgetting slows after the first year.
- At 11 months, participants "offered consistent answers about their flashbulb memories only 63% of the time".
- Over the next two years there was "a proportional decline of 9%, or an average of 4.5% a year".
- S12 against S13: t(390) = 5.21, p < .01, d = .28 (small).
- Emotional reactions are remembered worse than the other features.
- Overall consistency exceeded the consistency of the open-ended feeling: S12 t(375) = 8.72, d = .56; S13 t(364) = 9.17, d = .53 (both p < .01).
- The emotion-profile correlation fell from S12 to S13 (t(390) = 2.90, p < .05, d = .15). So "most of participants' forgetting of their emotional responses happened between Survey 1 and Survey 2".
- For one person, the correlation must exceed .73 to be significant at .05. Only 31.3% exceeded it at S12 and 25.3% at S13.
- Confidence stays high while consistency falls. The drop in confidence from Survey 2 to Survey 3 was not significant (p > .30).
- The retest check. Did filling in Survey 2 change later answers?
-
People who did only Surveys 1 and 2 (M = .63, SD = .20) were compared with people who did only Surveys 1 and 3 (M = .55, SD = .23): t(458) = 3.44, p < .01, d = .32.
-
The three-survey group's Survey 3 consistency did not differ from:
- the Surveys 1-and-3-only group;
- the Surveys 2-and-3-only group, scored with Survey 2 as the baseline.
Both ps > .40.
-
The samples did not differ in age, religion, residency, political viewpoint, gender or race/ethnicity (ps > .30).
-
Authors' conclusion: the pattern "is probably not a result of our retesting procedure".
-
Event memory (Table 5; mean accuracy, SD in brackets where given; the markers are those of the table, which does not name the comparison; the text implies change from the previous survey):
| Probe | Survey 1 | Survey 2 | Survey 3 |
|---|---|---|---|
| Number of planes | .94 | .86** | .81 |
| Airline names | .86 (.30) | .69 (.38)** | .57 (.42)** |
| Crash sites | .93 (.19) | .92 (.20) | .88 (.25)* |
| Order of events | .88 (.13) | .89 (.11) | .86 (.14)* |
| Location of Bush | .87 | .57** | .81** |
| … saw Moore's film | .87 | .60** | .91** |
| … did not see it | .86 | .54** | .71** |
| Overall | .88 (.14) | .77 (.21)** | .78 (.23) |
- Overall.
- Main effect of survey: F(1, 390) = 88.5, p < .001, ηp2 = .19.
- From Survey 1 to 2, accuracy fell by 13%: t(390) = 11.81, p < .001, d = .61.
- From Survey 2 to 3 there was no change: t(390) = .77, p = .45.
- The course depends on the fact.
- Crash sites and order did not change from Survey 1 to 2 (ps > .50).
- The number of planes declined only from Survey 1 to 2 (McNemar χ2(1) = 14.75, p < .001).
- Airline names declined continuously: Survey 1 vs 2 t(390) = 7.72, p < .001, d = .50; Survey 2 vs 3 t(390) = 6.21, p < .001, d = .30.
- Bush's location fell, then recovered.
- It fell from Survey 1 to 2 (χ2(1) = 95.43) and rose from Survey 2 to 3 (χ2(1) = 68.81), both p < .001.
- The authors call the rise the "Michael Moore Effect": the film Fahrenheit 911 showed Bush reading The Pet Goat in a Florida school.
- Viewers and non-viewers did not differ at Surveys 1 and 2, but did at Survey 3 (χ2(1) = 24.41, p < .001).
- Improvement was 52% for viewers against 32% for others. These are relative gains, .60 to .91 and .54 to .71 (our reading of Table 5).
- Consistency and accuracy were nearly unrelated. The only significant correlation was between S13 consistency and Survey 3 accuracy, r = .10, p < .05.
Predictors (Tables 6–8).
- Definitions.
- Residency: anyone living outside the city borders counts as a non-New Yorker.
- Living in downtown Manhattan made no difference (ps > .20).
- Where one learned of the attack had no effect.
- 40.4% of the three-survey participants reported a concrete personal loss or inconvenience. Psychological distress was not counted.
- Emotional intensity was taken from Survey 1, as the mean and as the maximum of the six ratings (r = .70 between the two).
- Flashbulb consistency: nothing predicted it.
- None of the five predictors (residency, loss or inconvenience, emotional intensity, media attention, conversation) related to consistency at S12 or S13. The consistency correlations in Table 7 lie between −.10 and .04.
- No single emotion correlated with consistency either.
- Non-New Yorkers were more consistent than New Yorkers: t(298) = 2.00, p < .05, d = .24. Table 6 marks the S12 difference (.65 vs .60). The authors are cautious because this misses their .01 level.
- Event accuracy: four of the five predicted it (all but emotion).
- In the ANOVA, residency F(1, 296) = 4.39, p < .05, ηp2 = .15; survey F(1, 296) = 38.54, p < .001, ηp2 = .12; loss × residency F(1, 296) = 7.22, p < .01, ηp2 = .15; survey × residency × loss F(1, 296) = 4.89, p < .05, ηp2 = .02.
- New Yorkers were more accurate: .81 vs .74 at Survey 2, .82 vs .75 at Survey 3 (Table 6; these New York cells are among those that do not fit their subgroups, see Limits).
- Among non-New Yorkers, those with a loss or inconvenience were more accurate: Survey 2 t(193) = 3.51, p < .001, d = .66 (.82 vs .70); Survey 3 t(193) = 2.65, p < .01, d = .44 (.82 vs .72). There was no such difference among New Yorkers (ps > .20).
- Correlations of accuracy with media attention and conversation (Table 7):
- Survey 1 accuracy with the first weeks' media attention (.19) and conversation (.16);
- Survey 2 accuracy with media .21 and .11, and conversation .19 and .18 (first weeks and first 11 months);
- Survey 3 accuracy with media .14 and .12 (first weeks, 35 months), and conversation .20, .11 and .12 (first weeks, 11 months, 35 months).
- Emotional intensity correlated −.01 to .07 with accuracy.
- Conversation mediates; media attention does not.
- Levels of media attention showed no effects (ps > .30).
- For the level of conversation: survey F(1, 285) = 265.34, p < .01, ηp2 = .48; residency × loss F(2, 285) = 3.77, p < .05, ηp2 = .04.
- Non-New Yorkers with a loss talked more at two weeks (t(192) = 1.97, p < .05, d = .28), 11 months (t(183) = 2.71, p < .01, d = .40) and 35 months (t(183) = 2.92, p < .03, d = .43). New Yorkers did not differ (p > .50).
- Mediation (Baron & Kenny, non-New Yorkers):
- none at two weeks, where the prerequisites failed;
- partial at 11 and 35 months (Sobel 2.01 and 2.04, both p < .05).
- Authors: "what mattered was not particularly where participants lived at the time of the attack or what personal loss/inconvenience they experienced, but how much they talked about the event."
- Confidence tracked rehearsal.
- Survey 2 confidence correlated with media attention (.13) and conversation (.18) over the first 11 months.
- Survey 3 confidence correlated with media attention (.28) and conversation (.20) over the 35 months.
How the content changed. The comparison is from Survey 2 to Survey 3 for flashbulb memory, and from each survey to the next for event memory.
- What stays consistent stays consistent, except emotion.
- The "objective" flashbulb features were place, informant, ongoing activity and the activity immediately following. For these, a consistent answer at 11 months was still consistent at 35 months 82% of the time.
- For the emotional reaction the figure was 44%.
- Wrong personal details are repeated.
- For objective features, an inconsistent answer at 11 months was more often repeated at 35 months than corrected or changed again (both p < .001).
- Authors: "a stable memory is forming after a year delay … even if they are full of inconsistencies with the initial report".
- Emotion reports keep changing. Participants more often reported a new emotion than they corrected or repeated a previous one (p < .01).
- Wrong facts are corrected.
- Event-memory errors were more often corrected than repeated or changed, at both later surveys. Corrections were more frequent than for flashbulb memories (all p < .005).
- Uncorrected errors: at Survey 2, a changed answer was as likely as a repeat. At Survey 3, repeats outnumbered changes, t(390) = 3.74, p < .001, d = .38.
- Authors: "Event memories are converging on an accurate rendering of the past, whereas flashbulb memories are converging on personally accepted and confidently held, even if inconsistent, renderings."
Discussion.
- Forgetting follows the ordinary curve.
- Forgetting slowed after the first year, as for ordinary autobiographical memory: "20% or more the first year and between 5% to 10% thereafter".
- Schmolck et al.'s acceleration may be an anomaly. Horn (2001) pointed to interference from the civil-trial verdict in month 16.
- Emotional reactions, despite being salient, "tend to be forgotten more quickly than other aspects of the flashbulb memory".
- Event memory follows media coverage.
- Figure 3 sets event-memory accuracy for 9/11 and for the Challenger explosion against the share of New York Times articles mentioning each event (both as z-scores). The pattern was the same with The Boston Globe and U.S. News and World Report.
- This is how the authors explain why forgetting did not slow in Bohannon and Symons's Challenger data.
- Memory practices.
- Media and conversation are a community's memory practices.
- Media retellings are fact-checked and act as "social/cultural reality monitoring". So event memories are corrected, except for details the retellings leave out, such as the airline names.
- Flashbulb memories are not retold by the community and cannot be checked, so their errors persist: "If a person falsely remembered that she was at work the first year, she tended to continue to remember (falsely) into the third year that she was at work."
- Why nothing predicts consistency. Factors may combine differently in different people; for example, some people rehearse an emotional memory and others avoid it.
Limits
- Consistency is not accuracy.
- Flashbulb answers are scored against the person's own report at one week. That report is itself a reconstruction, and there is no record of the circumstances.
- Consistency and accuracy barely correlated (r = .10 at best).
- Sample.
- A convenience sample: tables on campuses and in neighbourhoods, plus acquaintances of lab members. It is unweighted.
- 18% of Survey 1 respondents completed all three surveys, so the analysed sample is self-selected.
- Most analyses of residency have about 300 degrees of freedom (e.g. t(298), F(1, 296)) rather than 390. The text does not say why some participants were dropped.
- Group means only.
- There is no analysis of individual differences in forgetting, apart from the emotion-profile correlations. Whether a person's rate of change is stable was not studied.
- No per-person relation between confidence and consistency is reported.
- The forecast questions (expected accuracy in two or seven years) are described but never reported.
- Confidence data are thin.
- Only half of the Survey 2 and 3 participants were asked for confidence, and the n behind Table 4's confidence means is not given.
- Survey 1 apparently asked for no confidence: the confidence question is described as a change introduced in Survey 2, and Table 4 has no Survey 1 value (our reading).
- The emotion measure is ambiguous.
- Survey 1 asked about current feelings "at this moment", a week after the attack. The text treats these ratings as the initial reaction and the later ratings as memory of it.
- The wording in Surveys 2 and 3 is not given, and the surveys could not be obtained. If the later surveys also asked about current feelings, the z-score correlations measure how feelings changed, not how they were remembered.
- The open-ended feeling (Q4) is not affected by this.
- A part–whole comparison. Overall consistency includes the emotion feature. This makes the overall-versus-emotion contrast conservative.
- A weak retest check. It compares groups; the Surveys 1-and-3-only group has 72 people (Table 1); and it looks at mean consistency only.
- Illustration, not test. The media-coverage comparison (Figure 3) covers two events at three time points.
- Inconsistencies found:
-
Significance level. It is set at .01, but several effects presented as findings reach only p < .05:
- the residency main effect on accuracy, F(1, 296) = 4.39;
- the three-way interaction;
- both Sobel tests (2.01 and 2.04);
- the consistency–accuracy r = .10;
- several Table 7 correlations.
-
Effect-size labels contradict the stated benchmarks. d = .32 is called "medium", and the airline declines (d = .50 and .30) "large effect sizes".
-
ηp2 does not fit F for two effects. Residency F(1, 296) = 4.39 and loss × residency F(1, 296) = 7.22 are both given ηp2 = .15. F/(F + df2) gives about .015 and .024 (our computation). The other two effects in that ANOVA fit (.12, .02).
-
Degrees of freedom.
- Survey has three levels, yet its main effects are reported with one numerator degree of freedom: F(1, 390) = 88.5, F(1, 296) = 38.54 and F(1, 285) = 265.34. Survey is also called "dichotomous", and the text does not say which surveys enter.
- The 2 × 2 residency × loss interaction is reported as F(2, 285).
- The description of one ANOVA swaps the dependent and independent variables.
-
Table 6's New York columns cannot all be right. A group mean must lie between its two subgroups, but here it does not:
- accuracy at Survey 2: New York .81 against .84 and .89;
- accuracy at Survey 3: .82 against .75 and .75;
- consistency S12: .60 against .63 and .65;
- confidence S13: 4.32 against 4.11 and 4.27.
Two more problems in the same table:
- The S12 confidence subgroups are identical for New Yorkers and others (4.68 / 4.51).
- The full-sample S12 confidence (4.53 / 4.52) does not match Table 4's 4.41.
Table 6's other full-sample cells do match Tables 4 and 5.
-
Significance marks disagree between text and tables.
- The Survey 3 loss effect in non-New Yorkers is p < .01 in the text but marked p < .05 in Table 6.
- The two-week conversation difference is p < .05 in the text but marked p < .01 in Table 8.
-
The Discussion's forgetting figures fit only part of the data. "20% or more the first year and between 5% to 10% thereafter" fits flashbulb consistency (37% inconsistent at 11 months, our computation), not overall event accuracy, which fell 13% and then not at all.
-
Numbers of locations and return rates.
- The abstract says "seven US cities" and the Method "six recruitment locations"; Table 1 has seven columns.
- Table 1 gives (391 + 387) / 2,117 = 37% of Survey 1 respondents completing Survey 2, not the stated 38% (our computation).
-
Feature lists.
- The content analysis names "the activity immediately following the reception", but Table 2's question 6 asks what the person was doing "immediately before".
- Table 3's coding example ("Whom were you with …") is not among Table 2's questions.
-
Citations and references.
- Schmolck et al. is cited twice as 2002; the reference list has 2000.
- Some reference metadata are wrong. Luminet et al. 2004 carries the volume and pages of McCloskey et al. 1988. Davidson et al. is 2005 in the text and 2006 in the list.
-
What it means for Kurisutina
- Question 3: the size of the divergence the candidate prediction expects.
- For the personal circumstances of a real, emotional event, people's later reports agreed with their own one-week report on 63% of the six features at 11 months and 57% at 35 months (emotion included). For how they felt alone, the figures were 42% and 37%.
- A replica that keeps and retrieves its first account verbatim would score 1.0.
- So for this kind of content, perfect retention would separate the replica from its person on roughly a third of features within a year (1 − .63), and on more than half of emotion reports.
- This is group-level evidence for one kind of event: a public shock, mostly heard about rather than lived through, in a convenience sample.
- Current LLM memory systems are not verbatim either (venkit2026). So the replica's own consistency has to be measured, not assumed.
- The target shape is "change early, then settle", not continuous decay.
- Most change happens in the first year.
- After that, what is consistent stays so (82%), and personal details that were wrong at 11 months tend to be repeated at 35 months, while mean confidence stays high.
- A replica that keeps rewriting its memories would diverge from this in the other direction. The rate of rewriting is set by its reflection schedule (microverse2026).
- Reference measures for the planned test. Person and replica take in the same new material and are tested at
delays with free recall, lures, misinformation and reactivation. These measures carry over:
- Consistency across delays.
- Score each feature 1/0 against the person's own first report, with at least two later delays, so that the slowing can be seen.
- Human reference: .63 and .57 overall (six features, emotion included); .42 and .37 for the felt emotion; mean emotion-profile correlation .48 and .42.
- With controlled new material there is also a record, so accuracy can be scored beside consistency. Hirst shows the two can dissociate.
- Confidence against consistency.
- People held these memories at 4.41 and 4.25 on a 1–5 scale while about 40% of features had changed (our computation from 1 − .63 and 1 − .57).
- Confidence correlated with rehearsal (media attention .13 to .28, conversation .18 to .20). Its relation to consistency person by person is not reported; at group level it stayed high while consistency fell.
- Record confidence per item, and score the replica on whether its confidence relates to fidelity the way the person's does, not on calibration.
- The test must measure that relation itself (see also the confidence transfer in nichols2019).
- What decays and what is kept. Score each kind of content separately (as nichols2019 requires for kinds of
distortion):
- personal circumstances settle after the first year;
- reports of felt emotion keep changing, and new emotions appear;
- central facts change little (crash sites .93 → .88; order .88 → .86);
- peripheral facts nobody retells keep declining (airlines .86 → .57);
- re-publicised facts come back (Bush's location .57 → .81).
- Transitions between later tests. For each item wrong at the first delay, is it corrected, repeated or changed at the next? This is the finest-grained human reference here. It separates a person who settles on a version from a replica that either never changes or keeps changing.
- Post-event environment as a covariate.
- Fact accuracy followed conversation and media exposure, and conversation carried the effect of personal loss.
- Record what the person read, watched and talked about between tests, and give the replica the same. Otherwise a difference measures unequal inputs, not memory formation.
- Hirst's two 1–5 ratings per survey are the minimal version.
- Consistency across delays.
- No predictors of which details change.
- Residency, loss, emotional intensity, media and conversation did not predict flashbulb consistency.
- A replica can be scored on rates, transitions and profiles. It cannot be expected to predict from such factors which item a given person will change.
- Whether a person's own rate of change is stable was not studied; with nichols2019, it has to be measured in the person.
- Emotion is where a perfectly retaining replica should diverge most (our inference).
- The Q3 test should score reports of how the experience felt apart from facts and circumstances.
- It should ask the same wording at every delay (recalled feeling "then", not current feeling), to avoid Hirst's ambiguity.
- Provenance (brief 5.3, 6.4).
- Consistency with the first report tracks what is "held by the person" over time; accuracy against a record is "historically supported".
- A detail the person now reports confidently but that contradicts their first report is a case for the "contested" state of 5.3, with the first report as the record. A faithful replica keeps what the person now holds and keeps the first report as its source ("If the person remembers incorrectly, the copy inherits the error").
- Elicitation: interviews reactivate memories and may change them (brief 6.2, 6.3, 7).
- Each survey was a reactivation. An intermediate survey at 11 months did not detectably change consistency at 35 months (ps > .40). That is weak evidence: a between-group check, 72 people in the key group, mean scores only, and, as far as the text describes, no new material given with the survey.
- It does not show that interviewing is harmless. stjacques2013 shows that a matched cue followed by related new material changes memory within days.
- Hirst's design (groups that skip an intermediate survey) is the between-person form of the brief's perturbation audit with non-elicited controls. Kurisutina can hold back random items or delays from re-interviewing in the same way.
- The 6.2 rule "Keep first responses separate from later revisions" is what makes this kind of change measurable. Without the one-week report the change would be invisible, since the person is confident.
Cross-references
summaries/carry_on/stjacques2013_reactivation.md: reactivation with new material changes a real memory within days; the short-term, experimental complement to this study.summaries/carry_on/schacter2011_adaptive_distortion.md: why such changes happen (adaptive processes; Hupbach's reminder paradigm).summaries/carry_on/nichols2019_false_memory_tasks.md: distortion is stable within a kind, not across kinds; confidence transfers more than errors do.summaries/carry_on/venkit2026_companion_drift.mdandsummaries/carry_on/microverse2026_identity_drift.md: what LLM memory keeps and how often an agent rewrites itself.summaries/carry_on/park2023_generative_agents.md: memory stream, retrieval and reflection in an LLM agent.summaries/memory/johnson1993_source_monitoring.md: reality and source monitoring, which the authors invoke for media correction ("social/cultural reality monitoring").docs/research/elicitation_design_2026-09-22.mdanddocs/research/memory_retrieval_as_intervention_2026-09-22.md: earlier design notes on elicitation and on retrieval as an intervention.