Research date: 2026-09-22.
Citation and status. Dharshan Kumaran, Demis Hassabis, and James L. McClelland. What Learning Systems do Intelligent Agents Need? Complementary Learning Systems Theory Updated. Peer-reviewed Feature Review, Trends in Cognitive Sciences 20(7), 512–534, July 2016. DOI; PubMed. This is a theoretical review and integration of earlier modeling and experiments, not a new participant experiment or a systematic review with an explicit search protocol.
Acquisition and reading. Read the complete 23-page published-layout Stanford course copy: main text, Trends, Glossary, Boxes 1–10, model explanations, every figure/caption, Outstanding Questions, and bibliography. Visually inspected both main figures and all six box figures on printed pages 513, 515, 517, 518, 520, 523, 526, and 529. No separate experimental Methods section appears. The retained 44-page author manuscript was read only through its opening four pages for identification/version context, not independently in full; it explicitly labels itself a preprint and warns that citation numbering/details differ. No separate supplement was identified in either text or bounded primary-source searching; the publisher article endpoint returned HTTP 403, so its current complete attachments list was not inspected. No simulation rerun, data/code audit, or full reading of all cited studies was performed. Public Stanford copies were retrieved using verified TLS; publisher copyright remains, and no open redistribution license was identified. Published PDF, readable text, and source/version/hash manifest.
Source findings — factual core, under 200 words. The review retains complementary systems for structured knowledge and rapid episode-specific learning while revising their division of labor. Neocortical learning can be fast when new information fits existing representations; its effective speed depends on prior knowledge rather than anatomy alone. The REMERGE model allows recurrent retrieval of separate traces to support cross-episode inference. Encoding-based integration and retrieval-based inference remain incompletely distinguished empirically. Replay may interleave recent experiences with other memories or generated patterns, and its selection may prioritize significant experiences rather than reproduce exposure frequencies. The review relates these proposals to reinforcement-learning replay buffers and networks with external memory. It distinguishes stabilization within a memory system from integration across systems, and acknowledges that overnight behavioral improvements need not demonstrate rapid cortical integration. Open questions include maintenance of hippocampal indices as cortical representations change, replay-induced bias, and the conditions enabling rapid consolidation. These are computational proposals supported by selected earlier findings; the article neither reconstructs an unknown personal memory nor demonstrates transfer to another learner.
Critical audit — our assessment of what the review licenses.
- Comparison with 1995. The update makes prior-dependent speed and hippocampal inference central amendments (pp. 522–527). The original account already allowed variable consolidation times and some fast learning outside the hippocampus; it should not be rewritten as an absolute prohibition. The 2016 review also acknowledges that the original interference simulations used new information inconsistent with prior knowledge. A failure under that training regime is not a theorem about all sequential learners.
- A schema is not the new binding. The schema discussion (pp. 525–527; Box 8) compares category-feature simulations with the event-arena studies. In the already-read Tse experiment, new flavor–location relations were arbitrary bindings within a learned framework, and a training trial contained three rewarded collection trips. Do not turn “consistent” into “predictable from prior knowledge,” or “one trial” into one sensory exposure. The related 2011 gene-expression study and 2013 schema simulation are described by this review; their full originals were not newly read here. An analogy between their effects does not equate gene expression, model weight changes, and identified memory content.
- Inference does not locate a stored conclusion. REMERGE's account (Box 6; pp. 522–524) supplies a mechanism by which separately represented episodes can jointly support an answer. Therefore successful cross-episode inference alone cannot establish that the inferred relation was already stored as an integrated trace. The paper explicitly permits mixtures of encoding and retrieval mechanisms. A functional system can likewise answer from recomputation, an explicit stored conclusion, or both; correctness alone does not identify which operation occurred.
- Replay evidence has several endpoints. The broad consolidation language on p. 521 should not erase the limits of individual experiments. Our full Girardeau reading establishes a ripple-timed intervention effect on later spatial performance, without decoding content or measuring cortical information transfer. The review's reference 76, Bendor and Wilson, demonstrates auditory cue bias of spatial reactivation, without a later memory-benefit test. It does not itself isolate preferential replay caused by reward value. Neither result alone identifies the review's proposed goal-dependent weighting mechanism. Box 8's alternative of strengthened hippocampal traces is a useful warning against inferring systems transfer from improved behavior after sleep.
- Architecture claims need a stated regime. The opening assertion about agents needing two systems is stronger than a universal necessity result established here. Learning speed, interference, and capacity depend on representation, update rule, task distribution, available storage, and the tested resource budget. Fixed shared weights, an expandable store, isolated task parameters, and context-dependent retrieval impose different constraints. The review's neural-Turing-machine and memory-network comparisons (pp. 529–530) motivate functional separation of representation and rapidly changing records; they do not uniquely select a biological implementation or prove that a physical duplication of hippocampus and cortex is required.
Implications for extraction and transfer — original reasoning.
A stored code and its interpreter form a coupled system. The review's stable return mappings in Box 2 and unresolved index-maintenance question make an important transfer problem explicit: preserving an old code does not ensure preservation of its meaning after the recipient's representations change. A successful transfer criterion should test old records after recipient updates, with unfamiliar arbitrary bindings as well as familiar structured content. Rebuilding a compatible interpretation may be sufficient; copying the donor's cellular format is not demonstrated to be necessary.
Likewise, preserving event frequency and reproducing a person's future choices are distinct objectives. Selective replay can improve a task policy while biasing beliefs about how often an event occurred. It can also amplify an erroneous inferred detail if that detail is subsequently treated as an observed episode. A person model should retain evidence provenance and distinguish observations, inferred relations, and changes in retrieval priority. The review does not show that a person's replay priorities or learning rule can be recovered from an external memory archive.
Proposed discrimination, not a completed experiment. Hold the acquired episodes and information available to each recipient fixed. Cross recipient prior knowledge with three new-content classes: predictable consequences, fresh arbitrary bindings within a familiar domain, and explicit exceptions. Compare retrieval with a fixed interpreter, updating shared parameters, and a hybrid; specify storage and computation budgets rather than equating biological time with software time. Test immediate use, old-item interference, new combinations, and retention after the interpreter changes. Score actor decisions separately from forecasts of a source person's decisions.
For replay selection, compare exposure-proportional, significance-prioritized, and exception-preserving policies on the same records. Measure frequency calibration separately from task reward and source-detail fidelity. A policy that improves decisions while distorting event frequency should receive both results. Such tests can establish useful computational tradeoffs without identifying a biological mechanism, transferring subjective experience, or proving that an individual's personal learning trajectory was copied.
Subsequent primary-model readings — checkpoint7. The full 2012 REMERGE paper and Appendix and 2013 schema-learning simulations are now read separately. They qualify the review-level account: retrieval inference requires shared features and supplied bindings; rapid learning includes new identifiers bound to familiar representations, with explicit limits for novel conjunctions and interference. No original model code has been executed. The earlier read-scope statement above describes the review-stage work.