Kurisutina

Brain-inspired replay for continual learning with artificial neural networks

Nature Communications 11, 4069 (2020); received 19 February 2020, accepted 23 July 2020, published 13 August 2020. DOI 10.1038/s41467-020-17866-2; PMC7426273; CC BY 4.0. Peer reviewed (named reviewers: Raia Hadsell, German Parisi, Friedemann Zenke; a peer-review file exists and was not read). Baylor, Cambridge, UMass Amherst, Rice. Code: github.com/GMvandeVen/brain-inspired-replay, MIT licence (not read). Provenance: papers/base/vandeven2020_brain_inspired_replay.provenance.json.

What was read

  • Every line of the PMC XML converted with tools/pmc2txt.py (1,324 lines after folding): main text, Methods (task protocols, architectures, training, baselines, generative replay, distillation, the five modifications, the regularization methods with their equations, the generator measures), figure captions (Figs 1–9), availability statements and the 79 references.
  • Every line of the 9-page Supplementary Information (pdftotext -layout, 516 lines): the derivation of the latent regularization term for the mixture prior, the likelihood and reconstruction measures, and Supplementary Figs 1–4 (hyperparameter grid searches; generator measures).
  • Not read: the peer-review file; the figures as images. Most results are reported only as curves and bars, so few exact accuracies are available; values below are those printed in the text, or read from a supplementary axis where stated.

Question

Can replay prevent catastrophic forgetting in deep networks without storing past data, and can it scale beyond toy problems?

Methods

  • Three continual-learning scenarios (the authors' earlier taxonomy): Task-IL, where the task identity is given at test time; Domain-IL, where it is not but the output set stays the same; Class-IL, where new classes must be told apart from all earlier ones without being seen together.
  • Benchmarks. Split MNIST (5 tasks of 2 digits), permuted MNIST (100 tasks, each a new pixel permutation), split CIFAR-100 (10 tasks of 10 classes). Measure: average test accuracy over everything seen so far.
  • Compared methods. EWC and online EWC (penalize change to parameters by their Fisher information), synaptic intelligence (SI; importance from each parameter's contribution to the loss path), context-dependent gating (XdG; needs task identity), learning without forgetting (LwF; current inputs labelled by the old model), standard generative replay (GR; a separate VAE generates past-like inputs, labelled by the previous model), and baselines: plain fine-tuning ("None") and joint training on all data so far (upper bound). Hyperparameters by grid search.
  • Replay schedule. A fixed number of replayed samples per mini-batch (128 or 256), so cost does not grow with the number of past tasks; current and replay losses weighted 1/N and 1 − 1/N for N tasks so far.
  • Brain-inspired replay (BI-R): five modifications.
    1. Replay through feedback: the generator is merged into the classifier as its own feedback connections (a VAE with a classification head), so one network is trained.
    2. Conditional replay: a Gaussian-mixture prior with one mode per class, so specific classes can be recalled.
    3. Gating by internal context: a different random subset of decoder units is silenced per task or class; possible without task labels at test time because only the decoder is gated.
    4. Internal replay: replay hidden representations, not pixels. Requires early layers that barely change; the five convolutional layers were pre-trained on CIFAR-10 ("to simulate development") and frozen.
    5. Distillation: replayed samples labelled with soft targets at temperature 2.

Results

  • Scenario decides everything. On split MNIST in Task-IL every method prevented forgetting. In Class-IL, EWC and SI "dramatically failed" and only GR learned all digits. The grid searches show EWC and SI ending at ≈0.20 average accuracy in Class-IL split MNIST for every hyperparameter value tried (axis 0.200–0.204), about what remembering only the last two digits gives.
  • Replay is cheap and robust. One replayed sample per mini-batch of 128 current samples still beat every non-replay method in Class-IL. A VAE with only 10 hidden units produced poor samples yet only moderately reduced GR's accuracy. A control that re-initialized the network before each task showed why: retaining is much easier than relearning from scratch, and the weak replay was not enough for relearning.
  • Standard GR does not scale. On 100 permuted-MNIST tasks it declined rapidly after about 15 tasks. On split CIFAR-100 it failed even in Task-IL, worse than fine-tuning. In Class-IL on CIFAR-100, every existing method that does not store data forgot severely; only data-storing methods (iCaRL, experience replay) were acceptable.
  • Context from the earlier study (Masse et al. 2018, cited): after 100 permuted-MNIST tasks, SI ≈82%, online EWC ≈70%, XdG ≈61%, SI with XdG ≈95%, the last needing task identity at test time.
  • BI-R scales. It beat SI after 100 permuted-MNIST tasks without task labels, and BI-R with SI was better still. On CIFAR-100 it nearly removed forgetting in Task-IL and was best in Class-IL among methods that store no data, though "substantially under" joint training. With SI added: 0.904 ± 0.005 average accuracy on permuted MNIST and 0.344 ± 0.002 in Class-IL CIFAR-100; without replay-through-feedback, 0.892 ± 0.004 and 0.334 ± 0.002.
  • Ablations. Internal replay mattered most. The components were complementary: together they gained more than the sum of their separate effects. All except replay-through-feedback were necessary; that one mainly saves a separate model. Replayed samples improved in both quality and diversity (modified IS, FID and precision–recall).
  • Interpretation offered. Replay keeps memories in function space (input–output pairs to preserve); regularization keeps them in parameter space (which weights may move). Class-IL needs the former. The two are complementary, as cellular and systems consolidation are in the brain.

Limits

  • Image classification only; the class-conditional parts need labels, and internal replay needs a frozen pre-trained front end, which the authors say may limit learning of out-of-distribution inputs.
  • Only average accuracy was measured; forward and backward transfer were not.
  • Replay here has no temporal structure (no sequences), unlike hippocampal replay.
  • Few exact numbers in the text; the main comparisons are in figures not read as images.

What it means for Kurisutina

  • The slot's write path needs replay, not only weight protection. Integrating a new belief or memory into the person's parameters while keeping old ones distinguishable is closest to the Class-IL case, where Fisher-style protection alone (EWC, SI) failed. Interleaving old material, as the CLS literature requires, is the working ingredient; protection is a useful addition.
  • A little replay goes a long way. One old item per batch of new ones prevented most forgetting. Consolidation sessions can therefore be small, which matters at this project's scale.
  • Freeze the base, replay at the level the slot sees. Internal replay worked because the early layers were pre-trained and frozen. That is the project's own layout: a frozen base and a trainable slot. The episode store already holds the person's real episodes, so the project can use exact replay of stored episodes and needs generative replay only for material it cannot keep.
  • Conditional and context-gated recall (choose which class or task to replay) is the mechanism a logged, steerable replay scheduler would need.

Cross-references

  • summaries/base/kirkpatrick2017_ewc.md: the regularization method compared here.
  • summaries/base/spens2024_generative_consolidation.md: generative replay as a model of systems consolidation.
  • summaries/memory/mcclelland1995_complementary_learning.md: interleaving as the remedy for interference.
  • summaries/memory/bendor2012_replay_bias.md, rudoy2009_targeted_reactivation.md: cues steering replay in animals and people.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.