PNAS 113(29): 8171–8176 (2016), DOI 10.1073/pnas.1525569113, PMC4961177. Free through the PNAS open-access option. No
Creative Commons licence is stated on the PMC page, and the PMC XML carries only the front matter. Read from the
public PMC article page (the <article> element converted with pandoc). Princeton University and SUNY Albany.
Provenance: papers/carry_on/coman2016_mnemonic_convergence.provenance.json.
What was read
- All 265 lines of the converted article: significance statement, abstract, results, SI methods text (the Mantel test), discussion, materials and methods, and references.
- Figures 3 and 5 were downloaded and inspected: convergence, alignment by degree of separation, and reinforcement or suppression effects.
- Not read: Figures 1, 2, 4 and S1 (network diagrams, equations, item coding); only their captions or context. The separate SI file was not read, beyond the SI methods text included on the page.
Question
When people remember the same material together in pairs, over a network of conversations, do their memories converge? Does the network's structure matter? Do the individual effects of being reminded (reinforcement) or of hearing related items but not this one (retrieval-induced forgetting) drive the convergence?
Method
- 140 Princeton students in fourteen 10-person laboratory communities.
- Material. Everyone studied 16 items: 4 projects each of 4 fictional Peace Corps volunteers, making 4 categories of 4.
- Procedure.
- Individual recall.
- Three 150-second typed chat conversations with different partners, jointly recalling the material.
- Individual recall again.
- Network. Clustered (two subclusters, clustering coefficient .40) or non-clustered (.00): seven communities each, with the same number and order of conversations.
- Measures.
-
Pairwise similarity: the share of the 16 items that both remembered or both forgot. Convergence is its mean over pairs; alignment is its post − pre change.
-
Item reinforcement/suppression from conversation content:
- mentioned: +1;
- unmentioned but in a mentioned category: −1;
- unrelated: 0;
summed over each person's three conversations.
-
- Analysis. Network-level ANOVAs, mixed models by degree of separation, Mantel tests, and item-level closeness-centrality regressions.
Results
-
Memories converge. Mean convergence rose from pre- to post-conversation recall (F(1, 12) = 85.62). Read off Figure 3:
- clustered: about .55 to .61;
- non-clustered: about .57 to .66.
The larger rise in non-clustered networks was only marginal (p = .054).
-
Closer in the network, more alignment (Figure 3). In clustered networks, alignment fell with degree of separation, from about .10 at one step to about .04 at four or five steps (linear, p < .006). Only pairs 1–3 steps apart aligned significantly.
- Mantel tests: observed alignment matched each condition's own topology (r = .45 and .21), not the other's.
-
Mechanism: what is mentioned is kept, related-but-unmentioned is lost. Post − pre recall change by "pure" reinforcement/suppression score, read off Figure 5A:
Score Change +2 about +.17 +1 about +.11 0 (unrelated) about +.05 −1 about −.06 −2 about −.10 - Linear trend F(1, 9) = 70.95.
- Twice-reinforced and once- or twice-suppressed items differed from baseline; once-reinforced items did not (p = .35).
-
Community memory follows the conversation. Items reinforced across a community's 15 conversations became more central to its shared memory (by quartile, linear F(1, 12) = 160.20). The item-level regression was significant in 11 of 14 networks.
Limits
- Students, lab-created communities, fictional material with a category structure, and three short typed chats.
- Memory is scored only as remembered or not; no distortion or false memory is measured.
- One recall shortly after the conversations; nothing about how long convergence lasts.
- Network manipulation also changed diameter and path length. The authors acknowledge this confound.
- The pure reinforcement/suppression analysis covers 11 of 14 networks; 3 had no pure baseline items.
- Inconsistencies found: none in what could be checked.
- The df follow multivariate repeated-measures tests (e.g. F(4, 9) with 14 networks and five levels) and are consistent with the stated network counts.
- Figure values are my readings of the bar heights.
What it means for Kurisutina
- Q2/Q3: conversations reshape what the replica will remember, in a predictable direction.
- Whatever gets mentioned in a conversation is reinforced. Related things left unsaid are suppressed below the level of unrelated things. This happens to listeners as well as speakers (per the cited Cuc et al. 2007).
- A replica that talks with people will, if it is human-like, drift toward remembering what its interlocutors bring up. That is legitimate change: the person would do the same.
- The human effect size gives a reference. Three short conversations moved memory overlap between partners by about .06–.11 of the item set.
- Design implication (inferred). An LLM replica's memory store usually keeps everything or overwrites on
contradiction (Mem0). It has no retrieval-induced forgetting: unmentioned related memories do not weaken.
- If Q3 fidelity means human-like memory change, then reinforcement from mention and suppression of related, unmentioned items are candidates for the replica's memory dynamics.
- The rates should be the person's own, as with the fading affect bias and rehearsal.
- Test (proposal). Have the replica and the person each "remember together" with the same interlocutor (scripted
mentions) about a set of the person's real episodes.
- Then compare later free recall of mentioned, related-unmentioned and unrelated episodes between replica and person.
- A human-like replica shows reinforcement and retrieval-induced forgetting, with the same ordering, at about the person's magnitude.
- A store-everything replica shows neither. An overwrite store shows something else.
- Separating legitimate from illegitimate drift. Convergence with interlocutors is a normal human outcome, so drift
toward the interlocutor is not in itself a failure. This matches Wagner's point on audience tuning.
- The failure is when the replica converges faster or more fully than the person would, or converges on the base model's framing instead of the interlocutor's content.
Cross-references
summaries/carry_on/wagner2024_shared_reality_retrieval.md: audience tuning and memory.summaries/carry_on/hirst2009_911_memory.md: long-term change in flashbulb memories.summaries/carry_on/muir2022_neuroticism_fab.md: rehearsal keeps affect.summaries/carry_on/chhikara2025_mem0.md,summaries/carry_on/xu2025_amem.md: LLM memory stores without retrieval-induced forgetting.