McClelland, McNaughton, and O'Reilly 1995 — complementary learning systems
Reference: Why There Are Complementary Learning Systems in the Hippocampus and Neocortex: Insights From the Successes and Failures of Connectionist Models of Learning and Memory. Psychological Review 102(3), 419–457. DOI, author-hosted PDF. Theory and computational modeling, informed by prior empirical studies; no new human or animal experiment.
Reading record. Full main text now read, including model descriptions, simulation procedures, discussion, qualifications, footnotes, and all figure captions (printed pp. 419–453). This completes the earlier partial reading through extracted text line 1230. Read through line 2260, including the beginning of the references; the remaining bibliography was not read entry by entry, and cited studies were not all obtained. Equations 1–10 were visually checked in rendered PDF pages 437, 439, 445, and 447 where extraction was damaged. Table 1 was also visually inspected. Other graphical plots were not independently digitized or audited. No simulation rerun or source-code audit was performed. Source and hashes.
Source findings and limits, concise synthesis. The paper explains complementary fast episodic acquisition and slower structured learning through connectionist simulations. Interleaving protects existing responses while integrating new information; focused training can cause substantial interference. Consolidation simulations fit earlier amnesia data while treating the hippocampus as an assumed source of training examples, not an implemented retrieval network. A simplified strength model separates storage, decay, reinstatement, and response probability, with acknowledged parameter confounds. The binding discussion requires learned compression/decompression mappings beyond the stored hippocampal pattern. It considers fast learning outside the hippocampus and alternative anatomical boundaries, rather than insisting all cortex learns slowly. Statistical and gradient-learning arguments depend on the stated learning regime. Empirical timescales are fitted or interpreted, not independently predicted from cellular measurements. This is a computational rationale and family of models, not a demonstrated extraction procedure, unique biological account, or proof that every artificial learner must consolidate by changing weights.
Our interpretation: what would have to cross into a recipient?
A compressed representation is useful only with a compatible interpreter. Copying a code without the learned mappings that give it meaning does not follow from this account. Conversely, an artificial recipient may need a different code to preserve the same distinctions. The extraction target is therefore relational: what information must be supplied given the recipient's existing knowledge?
That question produces a concrete counterexample to evaluating extraction by narrative plausibility. A recipient with a strong prior could supply a coherent story while missing the arbitrary detail that distinguishes the actual episode. A weaker recipient could receive that detail correctly yet fail to use it. Neither case is adequately diagnosed by asking whether the generated story sounds like a memory. Tests should vary the recipient's prior knowledge separately from the transferred evidence, then assess both new details and their use.
A division between rapid storage and gradual integration is a useful hypothesis, but it does not select one software architecture. External records, a fixed base model, changing retrieval state, parameter updates, and combinations can preserve different functions. The relevant comparison holds acquired information constant and measures retention, source binding, exception handling, interference, generalization, and personal response fidelity separately. Biological timescales do not automatically specify artificial computation schedules.
An identifiability check — our derivation using the paper's simplified model.
The paper's equations 3, 5, and 6 use hippocampal strength S_h, cortical strength S_c, decay rates D_h and D_c, consolidation coefficient C, and retrieval factor R_h:
ΔS_h = −D_h S_h
ΔS_c = C S_h (1 − S_c) − D_c S_c
b_h = R_h S_h
Consider replacing S_h(0) with a S_h(0), C with C/a, and R_h with R_h/a, while holding D_h, D_c, S_c(0), R_c, and the baseline response probability b_p fixed. For any positive a that keeps parameters and states admissible, the entire S_h trajectory scales by a, while C S_h, S_c, and b_h stay unchanged. Therefore the model's resulting response probabilities also remain unchanged. At the factorized level, the paper defines C = ε r_h; choosing r_h → r_h/a with ε fixed also preserves reinstatement probability p(t) = r_h S_h(t). The same output equivalence holds if the hippocampal contribution is removed at a specified time. This is an exact observational ambiguity within this simplified model, not a claim about every biological experiment.
For example, initial strength 0.6, retrieval factor 0.8, and consolidation coefficient 0.1 can be changed to 0.72, 2/3, and 1/12 without changing those outputs. This makes precise why fitting performance alone need not recover a unique stored strength or consolidation coefficient. Independently informative measurements or interventions, plus their own measurement assumptions, would be needed to distinguish such parameterizations. Better fit from a personal parameter does not by itself make that parameter an identified personal mechanism.
Stability and change — a separate elementary derivation.
Consider a simpler learner, w_t = (1 − η)w_(t−1) + ηx_t, with fixed 0 < η ≤ 1. For independent observations from a fixed distribution of variance σ², its long-run variance is ησ² / (2 − η): solve V = (1 − η)²V + η²σ². A smaller rate reduces this variance. If the mean abruptly changes from μ_old to μ_new and the learner's expected starting value is μ_old, its expected error relative to the new mean after k observations is (1 − η)^k(μ_old − μ_new). A smaller rate prolongs the lag.
This is not a simulation or reanalysis of the paper's networks. It shows why a learning rate cannot be judged without specifying stability versus change in the environment. Faithful personal adaptation can also differ from statistically optimal estimation. A proposed personal update model should predict stable periods and controlled changes, and exceed a common update rule with matched estimates of initial beliefs and encoding. Retaining examples, retrieving exceptions, detecting context changes, or changing the learning rate are competing solutions.
Proposed experiments and next comparisons.
Cross familiar versus unfamiliar structure with arbitrary episode-specific bindings. Hold the extraction channel constant while varying what structured knowledge the recipient already has. Include answers that follow from the prior alone and answers that require newly acquired information. Freeze evaluation details before acquisition. Compare episode retrieval, replay-based integration, and their combination at matched evidence and stated computation budgets. Give every system the same human observations when testing architecture; retain unknown event truth only for evaluation when testing extraction.
A second manipulation should vary replay selection: complete sourced episodes, summaries, and mixtures containing exceptions. Test whether apparent compression benefits survive evaluation on source details and exceptions, and whether replay amplifies an earlier unsupported inference. This is a proposed artificial-system experiment, not a claim that biological replay literally rehearses cleaned verbal stories.
Read alongside the causal ripple study and schema-learning study in the next batch. Keep distinct: evidence that a process supports later performance; evidence about the content it carries; and evidence that this content is sufficient for a different recipient. Later theoretical revisions and contrary primary evidence remain necessary even though this original paper's main text is now read.