Nature Neuroscience 26, 1438–1448 (2023); received 24 October 2022, accepted 13 June 2023, published 20 July 2023.
DOI 10.1038/s41593-023-01382-9; PMC10400413; CC BY 4.0. Peer reviewed (one named reviewer: Kenneth Norman). Janelia
(HHMI), Harvard, Oxford, UCL. Code: github.com/neuroai/Go-CLS_v2, Zenodo 10.5281/zenodo.7941122 (not read). A
theoretical study: "We did not collect any experimental data". Provenance: papers/base/sun2023_go_cls.provenance.json.
Why this paper. It replaced Wu et al. 2022 (Memorizing Transformers) in the candidate list. A kNN memory added to a transformer is already represented in the library by EM-LLM, which builds on that line; this paper instead answers the question B3 depends on: which memories should consolidate.
What was read
Every line of the PMC XML converted with tools/pmc2txt.py (904 lines after folding): abstract, results, discussion,
Methods (architecture, training, retrograde amnesia curves, diverse teachers), figure captions (Figs 1–5), data and code
availability, the 67 references. Not read: the Supplementary Information (Supplementary Figs 1–7, the full
mathematical analysis in Sections 1–12, the deep-network checks on MNIST, CIFAR-10 and Tiny ImageNet in Supplementary
Fig. 5, the discussion of past experiments in Section 11 and proposed tests in Section 12), the Reporting Summary and the
figures as images.
Question
Why do only some hippocampal memories consolidate into neocortex, and when is consolidation harmful?
Model: teacher, student, notebook
- Teacher (the environment). A fixed linear network producing pairs y = w̄·x + ε. Predictability is the signal-to-noise ratio, SNR = σ²_w/σ²_ε (with σ²_w + σ²_ε = 1).
- Student (neocortex). A size-matched linear network (input dimension N = 100), weights starting at zero, trained by gradient descent on mean squared error.
- Notebook (hippocampus). A sparse Hopfield network of M = 2,000–5,000 binary units (sparsity a = 0.05). Each experience is bound one-shot by Hebbian learning to a random sparse index. Replay: random initialization settles into stored attractors (nine synchronous cycles); each memory was reactivated with near-uniform probability; well below capacity, recall was mostly perfect.
- Consolidation is the student learning from notebook reactivations: 100 reactivated patterns per epoch, 500–5,000 epochs, learning rate 0.005–0.1 (results qualitatively unchanged for rates ≤ 0.1). Generalization error is measured on typically 1,000 fresh teacher examples.
- Two regimes compared. Standard consolidation: replay without limit, which minimizes error on the stored memories. Go-CLS: stop consolidation at the point where further replay starts to harm generalization.
Results
-
Unregulated consolidation overfits (Fig. 2; N = 100, P = 100 memories, M = 2,000, learning rate 0.015):
- noiseless teacher (SNR = ∞): generalization error falls monotonically, and memory transfer is complete;
- SNR = 4: generalization first improves, then degrades as the student fits the noise replayed with each memory;
- SNR = 0.05: any consolidation harms generalization.
-
Go-CLS. For predictable teachers it consolidates everything; for less predictable ones it stops early. The student then generalizes near-optimally but memorizes the training memories only partly, while the notebook still recalls them exactly. Both generalization and memory improve with predictability.
-
How the stopping point can be found (a "supervisor"):
- hold out part of the notebook's memories as a validation set that never trains the student, and stop when its error rises; works best with small validation fractions (10–20%);
- estimate SNR by maximum likelihood from the input–output covariance of the stored examples;
- a heuristic: initial learning speed rises monotonically with SNR.
The authors do not expect the brain to set aside validation data literally; the estimates "rely on prior knowledge" that could be meta-learned.
-
Retrograde amnesia, simulated (score = (E0 − Et)/E0; the system uses whichever module is more accurate; a lesion removes the notebook and stops consolidation):
- standard theory: always a temporal gradient;
- Go-CLS: graded amnesia for predictable experiences and flat amnesia for unpredictable ones; prior consolidation of predictable experiences flattens the gradient, down to no amnesia at all, which the authors liken to schema learning (Tse 2007);
- Go-CLS predicts that memory transfer and generalization gain are correlated across memories.
-
Why two systems (Fig. 4): student plus notebook generalizes better than the student alone (online learning, each example seen once) and better than the notebook alone (nearest-neighbour recall, poor in high dimensions). The gain is largest when data are limited and moderately predictable, near P/N = 1. That is also where unregulated consolidation overfits worst, like double descent.
-
Unpredictability takes several forms that act alike: noise, a nonlinear teacher the linear student cannot express, and inputs the student cannot observe. Attention to a subset of inputs makes predictability depend on the learner's state, which "can differ across individuals" and "might partially underlie the individual variability in memory consolidation".
-
Predictions stated in the discussion. Predictability, not detail, frequency, feature overlap or salience, should decide consolidation. "Frequent misinformation should be consolidated less than rare gems from a wise source", unless brains use frequency as a heuristic for predictability. Extreme misregulation might relate to PTSD. Arbitrary but reliable facts ("Paris is the capital of France") count as fully predictable.
Limits
- Linear teacher and student, one scalar output, i.i.d. Gaussian inputs; the deep-network checks are in the unread supplement.
- No experimental data; links to amnesia findings are postdictions that "require plausible assumptions that may be wrong", and how to quantify predictability in real experiments is open.
- No biological mechanism for the regulator; replay selection, neuromodulators and attention are named as candidates.
- Supervised regression only; reinforcement learning and language models are left for future work.
What it means for Kurisutina
- A principled gate for B3. Consolidate an episode into the person's parameters only while it improves prediction of held-out episodes of the same person; stop when validation error on the person's other material rises. This is the same shape as the project's requirement that consolidation be accepted only if the person's existing answers survive, and it gives that requirement a stopping rule.
- What stays episodic. Unpredictable co-occurrences should remain in the episode store, retrievable but not written into the slot. That keeps one-off events, and one-off conversational pressure, out of the person's parameters.
- A tension to decide, not to hide. The normative gate would weigh "rare gems from a wise source" over frequent misinformation. People may instead treat repetition as evidence. A replica is meant to be the person, so its gate must reproduce the person's susceptibility at human size, not be optimal. Which heuristic people use is an empirical question this paper does not answer.
- The gate itself can be a person-level setting. Where attention sets predictability, the same experience consolidates differently in different people.
Cross-references
summaries/memory/mcclelland1995_complementary_learning.md,kumaran2016_complementary_learning.md,tse2007_schemas.md,mcclelland2013_schema_learning.md: the CLS and schema results this formalizes.summaries/base/spens2024_generative_consolidation.md: a generative teacher–student model of the same transfer.summaries/base/vandeven2020_brain_inspired_replay.md: replay against forgetting in deep networks.