Nature Human Behaviour 8, 526–543 (2024); received 30 May 2023, accepted 5 December 2023, published 19 January 2024.
DOI 10.1038/s41562-023-01799-z; PMC10963272; CC BY 4.0. Peer reviewed (one named reviewer: Gido van de Ven). UCL
Institute of Cognitive Neuroscience. Code: github.com/ellie-as/generative-memory (not read). Provenance:
papers/base/spens2024_generative_consolidation.provenance.json.
What was read
Every line of the PMC XML converted with tools/pmc2txt.py (1,383 lines after folding): abstract, introduction, model,
results, discussion, Methods, figure captions (Figs 1–7), data and code availability and the 148 references. Not
read: the Supplementary Information (supplementary results with two figures, the VAE architectures, the biological
reading of the Hopfield network), the Reporting Summary and the figures as images.
Question
Can one mechanism, hippocampal replay training a generative network, account for one-shot episodic encoding, gradual consolidation, semantic memory, imagination, relational inference and schema-based memory distortion together?
Model
- Hippocampus (the teacher). A modern Hopfield network (MHN) memorizes each event in one exposure. Inverse temperature β = 20, high enough that retrieval lands on stored memories; very few spurious attractors were seen. Random noise put into the network retrieves stored memories, sampled roughly at random: this is the model of replay.
- Neocortex (the student). A variational autoencoder (VAE) trained on the replayed patterns (teacher–student learning). Its latent variables are placed in entorhinal, medial prefrontal and anterolateral temporal cortex; decoding back to sensory experience runs via the hippocampal formation.
- Main simulation. 10,000 Shapes3D images stored in the MHN; 10,000 replayed samples, drawn once (not re-sampled per epoch), train the VAE. Latent size 20, learning rate 0.001, KL weight 1, mean-absolute-error reconstruction loss, AMSGrad, at most 50 epochs with early stopping. The VAEs were "not optimized for performance"; their use is illustrative.
- Extended model: what gets stored is gated by prediction error. During perception the VAE's reconstruction error is computed per element. Elements above a threshold are stored in the MHN as sensory features; the rest of the event is stored as conceptual features (here simply a copy of the VAE latents). Recall pattern-completes both, decodes the concepts into a schema-based prediction, and overwrites it with the stored sensory details.
- Recall test input: a noisy version of the stored item, with 10% of values set to zero.
- Not simulated: decay, deletion or capacity limits in the hippocampal store; any person-level differences; any sequence (events are single images, or bags of words in the DRM simulation). "In these simulations, the main cause of forgetting would be interference from new memories in the generative model."
Results (all simulations; no human data were fitted)
- Replay trains the student. After training on replayed samples, the VAE reconstructs items from partial input.
- Semantic memory as a by-product. A support-vector classifier reading the object's shape from the latents (200 training examples per epoch) improves as the VAE trains; across epochs, rs(48) = 0.997 [0.987, 1.000]. Latent codes allowed few-shot decoding better than sensory input or intermediate layers did (shown in the unread supplement).
- Imagination and inference. Sampling, interpolation and vector arithmetic in latent space generate new scenes and answer "what is to A as B is to C" (Fig. 3, images only).
- Schema distortion grows with reliance on the student. MNIST digits recalled by the VAE become more prototypical: within-class variation falls (paired t(7,839) = 60.523, Cohen's d = −0.684).
- Boundary extension and contraction. Atypically zoomed-out views are recalled with larger central objects and zoomed-in views with smaller ones, matching the direction of human data (Park et al.; the human measure was a closer/further judgement, the model's an object-size measure).
- The threshold trades detail for capacity. A lower prediction-error threshold stores more sensory features (rs(3) = −1) and gives lower reconstruction error (rs(3) = 1): less distortion, less efficiency. A higher threshold gives more prototypical recall. Distortion appears before consolidation, because concepts are stored from the start.
- Context at encoding distorts later recall (the Carmichael 1932 effect): an ambiguous shape encoded with "cube" or "sphere" is recalled toward that concept.
- False memory (DRM). A VAE trained on ROCStories word counts (vocabulary 4,206; 300 latents) as background knowledge; each list stored as a unique context term plus its latent code. Recall from the context term often, but not always, produces the lure, forgets some studied words, and adds intrusions ("vet" for the doctor list). Lure recall rises with list length, as in Robinson and Roediger 1997 (rs(10) = 0.998; 400 trials per point).
- Lesions. The latents support semantic recall without the hippocampus, and remote memories survive better (the retrograde gradient). Episodic re-experience needs the return path via the hippocampal formation (as in multiple trace theory). Anterograde amnesia follows from losing the one-shot store.
Limits
- Toy domains and qualitative fits. Shapes3D, MNIST and word counts; comparisons with human data are about direction, not size. The model is not fitted to any person or group.
- Single-dataset consolidation. The authors list lifelong consolidation as open: latent representations drift as the student learns, and new memories can catastrophically overwrite consolidated ones. They suggest generative replay from the student itself (citing van de Ven 2020) as a fix, untested here.
- Latent codes stored in the hippocampus go stale as the VAE changes, degrading recall. The authors argue the hippocampus more likely stores concepts derived from latents, which are more stable, "for example, in humans, language provides a set of relatively persistent concepts".
- Backpropagation; biologically implausible as the authors note. Replay selection (reward, salience) not modelled.
What it means for Kurisutina
- A concrete B2–B3 pattern. An append-only, one-shot episode store (the teacher), replay sampled from it, and a slow parametric student trained on the replayed episodes. The student's prediction error decides what the store must keep in detail. This is the complementary-learning design with a working gate on encoding.
- The encoding threshold is a natural person-level setting. The authors propose that salience lowers it (traumatic memories in more sensory detail) and that weaker priors behave like a lower threshold. That gives B4 a route into memory: affect sets how much detail is kept.
- Distortion is a feature to reproduce, not a bug to remove. Gist-based intrusions, boundary extension and context-driven recall are human signatures that a correct base should show; they belong in the B9 battery.
- Store episodes in a base-version-independent form. Stored latent codes go stale when the student changes. The project's rule that slot codes belong to the base version that wrote them has a direct counterpart here: keep the episode store in text or concept form so it can re-train a new base.
Cross-references
summaries/memory/mcclelland1995_complementary_learning.md,kumaran2016_complementary_learning.md: the CLS framework this updates.summaries/memory/roediger1995_false_memories.md: the DRM effect simulated here.summaries/memory/tompary2017_consolidation_overlap.md: consolidation fading detail and growing overlap in people.summaries/base/vandeven2020_brain_inspired_replay.md: the replay method proposed for lifelong consolidation.summaries/base/sun2023_go_cls.md: when consolidation should and should not happen.