PNAS 114(13), 3521–3526 (2017); received 19 July 2016, approved 13 February 2017, edited by James L. McClelland.
DOI 10.1073/pnas.1611835114; PMC5380101; "freely available online through the PNAS open access option". DeepMind and
Imperial College London (14 authors, Kirkpatrick first, Hadsell last). Peer reviewed. A published comment exists
(Huszár, "Note on the quadratic penalties in elastic weight consolidation", PNAS 115, E2496) with the authors' reply
(E2498); neither was read. Provenance: papers/base/kirkpatrick2017_ewc.provenance.json.
What was read
- Published version: every line of the PMC article page converted to text (941 lines after folding): significance, abstract, main text, Materials and Methods, the supporting-information sections that PMC shows inline (Fisher overlap; the random-patterns derivation, Eqs. S1–S33; MNIST and Atari methods; Tables S1 and S2) and all figure captions (Figs 1–4, S1–S3), plus the references. The first conversion dropped the captions; the page was re-converted with captions kept and the captions read.
- Preprint: every line of arXiv 1612.00796v2 (13 pages, 767 lines), which has the same core text and appendix but not the random-patterns analysis or the cascade-model discussion.
- Not read: the figures as images (results are given as curves without printed numbers), the Huszár comment and reply.
- Internal inconsistencies noticed: the Atari methods text says replay buffer 5 × 10⁵ and target update every 3 × 10⁴ steps; Table S2 says memory 50,000 and target update every 7,500 steps.
Question
Can a network learn tasks in sequence without forgetting earlier ones, using one network of fixed size and no stored data, by protecting the weights that earlier tasks rely on?
Method
- Penalty. While learning task B, minimize L(θ) = L_B(θ) + Σᵢ (λ/2) Fᵢ (θᵢ − θ*_A,i)². Fᵢ is the diagonal of the Fisher information at the task-A solution: a Laplace approximation to the posterior over parameters after task A, used as the prior for task B. λ sets how much the old task matters. With more tasks, the penalties add.
- Biological framing. Synaptic consolidation: synapses important for a learned skill become less plastic (dendritic-spine persistence and erasure studies in mice are cited).
- Experiments.
- Random patterns: a linear network (n = 1,000 synapses) associating random binary patterns, analysed exactly and simulated (400 runs). A memory counts as retained when its SNR exceeds 1.
- Permuted MNIST: fully connected ReLU networks (Table S1: 2–6 hidden layers of 100–2,000 units, 20–100 epochs per dataset); compared with plain SGD, uniform L2, and SGD with dropout and early stopping (50 random hyperparameter settings per experiment).
- Atari 2600: a Double-DQN agent learning 10 games drawn from a pool of 19 that DQN plays at human level; 10 sets of 10 games × 4 seeds. Extras: unsupervised task recognition (a forget-me-not Dirichlet model, nonparametric, adds a new task when a held-out model wins), a short-term replay buffer per inferred task, game-specific biases and multiplicative gains in each layer, Fisher multiplier 400, EWC applied only after 20 million frames in a game.
Results
- Random patterns. With plain gradient descent, a pattern's SNR decays as a power law (slope −0.5) only until the number of patterns approaches capacity (t ≈ n), then exponentially. EWC keeps a power law throughout and retains a larger fraction of memories at capacity. Lower fixed learning rates delay but do not prevent the exponential decay. Past capacity, EWC does worse than gradient descent, and like a Hopfield network it can suffer "blackout catastrophe". Under EWC "weights can only become more constrained (i.e., less plastic) with time", so it models retention, not forgetting.
- Permuted MNIST. SGD forgets task A once training moves to B. Uniform L2 keeps A but cannot learn B. EWC learns B while keeping A, and scales to many tasks "with only modest growth in the error rates", where dropout with early stopping does not. (Results are curves; no numbers are printed.)
- Shared or separate weights. Fisher overlap between two tasks is high throughout the network for similar tasks (8 × 8 pixels permuted) and lower in early layers for dissimilar ones (26 × 26), while layers near the output stay shared.
- Atari. With SGD the agent "never learns to play more than one game", and the summed human-normalized score (max 10) stays below 1. With EWC it learns several games, but stays below 10 separately trained DQNs. Giving the true task label instead of the inferred one helped only modestly.
- The Fisher underestimates uncertainty. In a single-game agent, perturbations shaped by the inverse Fisher hurt less than uniform ones, but perturbing the Fisher's null space hurt as much as perturbing along the inverse Fisher, so some "unimportant" parameters matter. The authors call this the chief limitation.
Limits
- Diagonal Laplace approximation from a point estimate; acknowledged as a "significant weakness".
- Plasticity only ratchets down; no mechanism to recover plasticity or forget deliberately.
- Needs task boundaries (given, or inferred by a separate module) to know when to compute and anchor a penalty.
- Task-incremental style evaluation. Van de Ven et al. 2020 later showed EWC failing when new classes must be told apart from old ones (class-incremental learning).
- The Atari system mixes EWC with per-task replay buffers and per-game parameters, so the share due to EWC alone is not isolated.
What it means for Kurisutina
- A cheap guard for the slot, not a write path. A Fisher-weighted penalty can anchor the person's parameters to their previous state during each consolidation round, at a cost linear in parameters. It protects what the slot already holds; it does not integrate new material with old (that needs replay; see van de Ven 2020).
- Lifetime plasticity is the risk. A replica that carries on for decades must keep changing at a human rate. EWC alone makes the slot steadily stiffer and, past capacity, worse than plain training. Any penalty would need decay (online EWC's γ < 1) or a mechanism like the cascade models' that lets synapses become plastic again.
- Small context-specific parameters help. The agent's per-game biases and gains are a precedent for keeping session or context channels separate from the person core, as the architecture already requires.
- Importance estimates are overconfident. A gate that trusts the Fisher to say which slot parameters "do not matter" will let some person content erode; tests must measure retention directly, not infer it from the penalty.
Cross-references
summaries/base/vandeven2020_brain_inspired_replay.md: EWC fails in class-incremental learning; replay works.summaries/memory/mcclelland1995_complementary_learning.md: the systems-level alternative (interleaved replay) that EWC was designed to avoid storing data for.