Kurisutina

Haehner, Kritzler, Fassbender & Luhmann (2021/2022). Stability and change of perceived characteristics of major life events

Journal of Personality and Social Psychology 122(6), 1098–1116 (2022; online 30 September 2021 according to OpenAlex), DOI 10.1037/pspp0000394. The published version is closed access (OpenAlex finds no open copy) and was not read. Read as the PsyArXiv preprint, version 1 (the only version), in the third and latest revision of its file, uploaded on 28 May 2021. It states that it was accepted on 24 May 2021 and is "not the copy of record", so it is the accepted manuscript after peer review (DOI 10.31234/osf.io/2yzcs). The earlier revisions of 27 January and 19 May 2021 were not read. OSF lists it as CC BY 4.0, but the PDF carries an APA copyright line and asks readers not to copy or cite it without permission. Ruhr University Bochum and University of Siegen. Provenance: papers/carry_on/haehner_perception_stability.provenance.json.

What was read

All 2,761 lines of pdftotext -layout output (57 pages; my count with wc and pdfinfo): the title page with the acceptance statement, author note, abstract, introduction with Table 1 (earlier retest studies), method, results, discussion, references, Tables 2–4 and the captions of Figures 1–4.

Not read or not inspected:

  • the figures (model diagram, 75 individual trajectories, autoregressive curves, mean-level curves), which are images;
  • the supplemental material: Tables S1–S9, with the preregistration deviations, descriptives, event categories, measurement invariance and the other Big Five traits, and Figures S1–S4, with the model equations;
  • the preregistrations (osf.io/kvf5g; osf.io/urqdw, view-only link), data and scripts, and the published version.

Questions

  1. How stable is the rank order of a person's appraisal of the same event over a year, compared with Big Five traits and affective well-being?
  2. Do average appraisals drift? The hypotheses: extraordinariness falls, predictability rises, valence rises.
  3. What share of the variance is between persons (intraclass correlation, ICC)?

Method

  • Sample. The What's NEXT? Study, as in Luhmann et al. (2021) Study 5 and Haehner et al. (2022; haehner2022 summary). It uses a different event, though: the one named at T1.
    • 857 registered. After quality checks and a 15-week rule: N = 619, 430, 364, 331 and 321 at T1–T5. Mean age 21.48; at T1, 72.54% were female and 92.57% fall under Table 2's high-school graduation column.
    • Mostly positive events: vacation (52), starting college (48), relocation (47), Abitur (46).
  • Design.
    • Five online surveys, planned at 0, 12, 24, 36 and 48 weeks. Actual gaps averaged about 12–13 weeks (range 4.06 to 21.06).
    • At T1 people named the most important event of the last three months, on average 6.78 weeks earlier (range 0–15), and rated it on the 37-item ECQ version of Luhmann et al.'s Study 5.
    • They re-rated the same event at T2 and T5. At T3 and T4 only a random subsample did, so the design has planned missingness. The number of re-ratings at T3 and T4 is not given in the main text.
  • Benchmarks. Big Five (BFI-2-XS, 3 items each; only openness is reported in the main text; agreeableness was dropped for lack of measurement invariance) and affective well-being (6-item SPANE, last month).
  • Models. Continuous-time models (ctsem), one latent process per scale, with time zero at the event.
    • Single items as indicators; three two-item parcels for valence and affective well-being.
    • Rank-order stability: the continuous-time auto-effect a, converted into one- and twelve-month autoregressive coefficients.
    • Mean change: ES15, the change in the latent mean from the event to 15 months, in SD units at the event. It counts as significant if ΔAIC > 4 and the likelihood-ratio test is significant.
    • ICC: the long-range between-person share of variance, from random-intercept models.
    • Impact reached invariance only after one item was dropped.

Results

  • Rank order.

    • The auto-effect was negative for 8 of 9 subscales, so stability falls with the interval (Hypothesis 1 supported).

    • One-month coefficients: .961–.981 (impact aside).

    • Twelve-month coefficients, in ascending order:

      Measure Twelve-month coefficient
      Social status change .622
      Change in world views .634
      Challenge .698
      Predictability .711
      Valence .755
      External control .763
      Emotional significance .783
      Extraordinariness .793

      Benchmarks: openness .977, affective well-being .185.

    • Impact had a small positive auto-effect (0.010). It was not robust with the four-item scale (0.006, interval including zero).

  • Mean drift over 15 months (ES15, Table 4).

    • Significant by the joint rule: change in world views +0.41 (Δ−2LL = 28.19, p < .001) and extraordinariness −0.19 (8.08, p = .004).
    • Significant by the likelihood-ratio test alone, but ΔAIC < 4: external control +0.17 (p = .021), emotional significance −0.15 (p = .024), impact −0.14 (p = .033) and valence +0.10 (p = .018).
    • Not significant: predictability −0.08, social status change −0.11, challenge +0.06.
    • Benchmarks: openness < 0.01, affective well-being 0.07.
  • ICC, the long-range between-person share.

    Measure ICC
    Impact .94
    Valence .87
    Predictability .85
    Emotional significance .83
    Challenge .80
    External control .78
    Extraordinariness .76
    Social status change .68
    Change in world views .63

    Benchmarks: openness .98, affective well-being .53.

  • Person and event are not separated. The authors say so directly: "the participants experienced and rated different major life events, which obviously contributes to interindividual differences". A trait-like style of perceiving events is the other possible source.

  • The authors' reading. Reappraisal was limited. The rise in world-view change fits meaning-making taking at least a year. Hindsight and "rosy view" biases may already have been over by T1, about seven weeks after the event. They also write that "the high stability of the life event characteristics might imply that episodic memory of major life events is in general quite accurate over time".

Limits

  • Sample and events. Young, educated, mostly female and German. Events were mostly positive, so valence had little room to rise (the authors say so).
  • The early weeks are not observed. The first rating came 0–15 weeks after the event, and ES15 is measured from the event itself (time zero), before any rating. Part of each drift is therefore model extrapolation (inferred).
  • One event per person. The ICCs and rank-order coefficients mix a person's style with the particular event (the authors say so).
  • Stability is latent. The coefficients are free of measurement error, so the agreement of raw ratings over a year will be lower (inferred).
  • Strict significance rule. With 1 df, "ΔAIC > 4" means Δ−2LL > 6, i.e. p < .0143 (my computation). That is stricter than the α = .05 the paper states.
  • No link to outcomes. The study did not test whether re-appraisal goes with changes in well-being or personality (the authors say so).
  • Version. The earlier ECQ version (37 items), not the final 38-item one.
  • Inconsistencies found:
    • Table 4, social status change. Δ−2LL = 2.51 does not fit the printed p = .142 or ΔAIC = 0.15; both fit 2.15. This is verified by script; that 2.51 is a transposition of 2.15 is inferred.
    • Stated α against the decision rule. The paper states α = .05, but under it valence rose significantly (+0.10, p = .018) and Hypothesis 4 would be supported. The text says valence did not "significantly" change, and calls the changes in impact, emotional significance and external control "non-significant" at p = .033, .024 and .021.
    • ICC ranking. The text names impact, valence and emotional significance (.83) as the characteristics where between-person differences matter most. Predictability (.85) is higher than emotional significance.
    • Size of the drifts. The drifts are said to be "two to four times larger" than those of the Big Five and affective well-being. Against the Table 4 benchmarks (openness < 0.01, affective well-being 0.07), the range 0.06 to 0.41 is not 2–4 times. The other traits are in the supplement, which was not read.
    • Stability of valence and emotional significance. The discussion calls these stabilities "particularly high". Their twelve-month rank-order coefficients (.755, .783) are mid-range, and both had likelihood-ratio-significant drifts at .05.
    • Blinding. Placeholders ("[INSTITUTION BLINDED FOR PEER REVIEW]") remain in the accepted version, though the author note gives the links.
    • Arithmetic checked by script:
      • Table 3: the one- and twelve-month coefficients equal exp(a) and exp(12a) in all 11 rows.
      • Table 4: ΔAIC = Δ−2LL − 2 and the p-values hold in 10 of 11 rows.
      • Table 1: the error-adjusted stabilities equal r/√(α₁α₂), up to one rounding (Frazier, future control, 4–6 weeks: .798 printed as .79).

What it means for Kurisutina

  • Question 3: a person's appraisal of a remembered event is fairly stable but drifts in known directions.
    • Over a year the rank order holds at .622–.793 (latent). The means barely move except for two dimensions: events come to seem more world-view changing (+0.41 SD over 15 months) and less extraordinary (−0.19 SD). These numbers are verified; the reading as meaning-making is the authors'.
    • A replica that stores its first appraisal and never revisits it would miss that drift. One that rewrites freely would lose the rank-order stability. The revision rate of its reflection step should be calibrated against these values, as the goal doc already proposes for Hout's β (inferred).
    • This does not contradict hirst2009, where people's reports of how they had felt matched their own first report only 42% of the time at 11 months (hirst2009 summary). The measures differ: the rank order of a rating scale here, item-by-item consistency of content there. Stable ratings do not show accurate memory, contrary to the authors' remark (inferred).
  • Question 1: the 63–94% between-person share is an upper bound for a personal appraisal style, because every person rated a different event (verified).
    • With Kritzler's rough split (valence mostly set by event type, predictability mostly not; kritzler2022 summary), the person-specific part is unknown but plausibly largest for predictability-like dimensions (inferred).
    • Only a design with several events per person, or events shared across people, can separate the two (inferred).
  • Held-out appraisal test (goal doc, latest entry): concrete consequences.
    • Ceiling. The target is her rating at a given time. Measure her own re-rating agreement by having each held-out event rated twice, weeks apart. Report the replica's agreement beside that ceiling, as Hout's reliabilities are used for the GSS pilot (inferred proposal).
    • Time since the event. Record it for every rating. Give it to the replica and match the typical-profile baseline on it, because world-view change and extraordinariness drift with it (inferred).
    • An additional baseline: the person's own mean appraisal across her other events. If a stable personal style exists, that baseline beats the typical profile. An own-slot replica must then beat both (inferred proposal).
  • Measurement bursts (goal doc proposal 2).
    • In this sample, affective well-being had a one-year coefficient of .185 and an ICC of .53: about half its long-range variance is within the person. That supports separating current state from stable disposition (verified numbers; link to sliwinski2009 inferred).
    • A cheap addition: re-rate one or two reference events at every burst. That measures re-appraisal, and whether it tracks current state, at almost no cost (inferred proposal).
    • Continuous-time models handle unequal gaps and planned missingness, which suits bursts months apart (inferred).
  • LISS T1 and the GSS pilot. Neither records appraisal, so nothing here changes them.
    • The one link: a person's appraisal of one event seems to stay fairly stable within a year. Differences between two reactions are then more likely to come from appraising the two events differently than from re-appraising one of them (inferred).
    • Whether that holds over the GSS's two years is not known (guess).

Cross-references

  • summaries/carry_on/luhmann2021_ecq_taxonomy.md: the ECQ; perceived valence and SWB change in the same study.
  • summaries/carry_on/haehner2022_event_perception.md, summaries/carry_on/kritzler2022_event_perception_profiles.md, summaries/carry_on/rakhshani2021_traits_event_perception.md: the rest of the appraisal thread.
  • summaries/carry_on/hirst2009_911_memory.md, summaries/carry_on/schacter2011_adaptive_distortion.md: how memories of events change.
  • summaries/carry_on/microverse2026_identity_drift.md: revision rate set by a reflection schedule.
  • summaries/carry_on/sliwinski2009_stress_bursts.md, summaries/carry_on/hout2016_gss_reliability.md: bursts; reliability ceilings.
  • docs/research/liss_q1_design.md, docs/research/gss_pilot_design.md.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.