Kurisutina

Prior knowledge changes effective learning speed

Identity and status. James L. McClelland, Incorporating Rapid Neocortical Learning of New Schema-Consistent Information Into Complementary Learning Systems Theory. Journal of Experimental Psychology: General 142(4), 1190–1210. DOI: 10.1037/a0033812; PubMed metadata; Stanford author PDF. Published theoretical review with original computational simulations, not a new human or animal experiment. Accessed 22 September 2026. The retained PDF is the August 26 Online First version, with article pages 1–21; references below use those pages. Final paginated-version equivalence was not checked.

Factual core

A semantic network pretrained on eight items learns new identifiers associated with familiar feature combinations faster than an exception combining incompatible category properties. Its learning-rate parameter remains .1. Pretraining uses 1,000 epochs; subsequent focused/interleaved epochs contain 2/34 examples. New identifiers receive dedicated localist inputs. Consistent additions cause small, rather than universally zero, interference. Further simulations examine weight changes, gradient coherence, unfamiliar internally structured environments, and restricted plasticity. A competitive network illustrates familiarity-dependent activation and learning, with eventual replacement of old structure after an environment switch. These are qualitative computational accounts; animal findings are reviewed rather than newly measured. The discussion explicitly questions extension from remapping familiar representations to learning novel conjunctions. (Methods, results, and discussion, pp. 5–19.)

Printed inconsistency: Figure 6's caption says sum, but methods and axis say average squared weight change; its caption also reverses the dark/white condition labels relative to the legend. Interpret using methods and legend, pending code verification.

Original analysis: what this does and does not identify

Arbitrary binding and arbitrary content are different demands. Consider a new random identifier assigned uniformly to one of K familiar categories. Before the assignment is observed, an otherwise knowledgeable recipient still lacks log2(K) bits. Supplying that assignment can convey real new information even when all category properties were already known. Conversely, correctly supplying those properties afterward is not evidence that each was independently extracted. An extraction benchmark should score the unpredictable assignment separately from prior-supported consequences. It should then ask whether the same method handles an exception, a new combination of familiar attributes, and a genuinely new attribute. These are distinct tests, not progressively more verbose prompts.

Capacity allocation can make apparent integration cheap. In our abstraction, write the recipient as f_phi(e_x, r), with shared decoder parameters phi and a separate representation e_x for each identifier. If phi is frozen, updating only e_new cannot alter outputs for another fixed identifier. That protection follows from parameter separation; it does not establish a general solution to interference. Conversely, changing shared parameters can affect old outputs even when the new item is described as consistent. Compare a private writable slot, a distributed shared representation, and an external episodic lookup at matched capacity. Measure the whole old test set, not just properties closely related to the new item. Sparse distributed inputs may approximate isolation, but the required overlap and available capacity must be measured rather than assumed.

Coherent change is not a count of recovered facts. For update vectors d_t, the triangle inequality gives ||sum_t d_t||_1 / sum_t ||d_t||_1 between zero and one when the denominator is positive. This ratio distinguishes cancellation from aligned movement; neither numerator nor ratio measures information about a particular person. A large change can implement a generic rule, and many canceling updates can conceal substantial intermediate plasticity. Mapping either endpoint displacement or cumulative path length onto a biological gene-expression assay requires its own observation model. Their qualitative resemblance cannot identify synaptic storage or an exact replay count.

Speed needs two budgets. Equal epochs need not mean equal computation or equal exposure to old material. Report new-example presentations, all updates, replay volume, and elapsed computational cost separately. Test both equal new-exposure and equal compute budgets; these answer different questions. A reduction in old-error growth bought through additional old-example replay is a legitimate result, but its cost must remain visible. Simulation epochs cannot be converted directly into human encounters, hours of consolidation, or interview duration.

Representational consistency is not truth. For a personal model, an unusual but well-supported event can conflict with the recipient's prior. Rejecting or weakening it because it fits poorly could preserve a stereotype. Use separate variables for prior compatibility, documentary support, and the person's belief. The same message can be easy for one recipient and difficult for another because their representations differ. A helpful falsifiable claim concerns that interaction, not a universal ranking of records by intrinsic schema consistency.

A constructive account does not prove architectural necessity. These demonstrations make a fixed-learning-rate network a useful counterexample to an unconditional slow-learning premise. They do not exhaust possible learners. Claims that every successful recipient needs biological hippocampus/cortex analogues, or that all inconsistent information must be learned slowly, would require assumptions about capacity, architecture, training, and admissible interference that this paper does not establish as a general theorem. Nor does delayed replacement of an old representation establish its permanent preservation.

Falsifiable transfer experiment — proposed, not performed

Create controlled, evaluator-recorded episodes that assign randomly generated identifiers to categories or locations. Cross recipient prior training with message type: familiar-template binding, exception, new conjunction, and new feature. Counterbalance which mapping is familiar across recipients while holding the incoming message fixed. Include a no-record recipient and an equally informed exact-record lookup; keep the evaluator's assignment key out of ordinary recipient inputs. Prespecify prior-only guesses and isolate errors on the unpredictable assignments.

Compare frozen-decoder writable representations, shared-parameter updating, replay-assisted updating, and episodic retrieval under explicit memory and update budgets. Before evaluating, reserve old facts, new paraphrases, untrained relations, and reversal cases. Score calibration and assignment accuracy alongside interference and downstream choices that actually require the new binding. If gains occur only on inferable properties while random assignments remain at chance, the result supports useful prior completion without establishing acquisition of episode-specific information. If private slots match a shared learner, successful transfer does not establish that shared knowledge was reorganized. If the recipient takes correct new actions yet fails to forecast the source person's errors, functional usefulness and personal fidelity have diverged and must be reported separately.

Reading and archive record

Read the entire retained main paper: introduction, all methods/results/discussion for Simulations 1, 2, 2′, and 3, equations, all eight footnotes, nine figure captions, General Discussion, Conclusion, and references. All figures were visually checked in their surrounding pages; PDF pages 2, 4, 6–10, 12, and 14–17 were inspected, including architecture, training patterns, error axes, weight-change legends, and competitive-learning equations. Figure 3 uses a reversed error axis: higher on the page means less error. No numerical graph digitization or replication was performed.

No supplementary document is mentioned in the article or linked beside this paper in the verified Stanford publication entry. The cited PDPTool handbook/software was not downloaded or executed; it is an unread implementation boundary. The publisher DOI page was inaccessible through the browser, so absence of a publisher-only supplement is not independently guaranteed. This reading does not newly count cited animal experiments as completed primary readings. The author-hosted copy retains APA copyright; no open reuse license was identified.

The original 22-page PDF is retained with a readable text extraction. Standard pdftotext preserves page breaks but sometimes presents two-column text in awkward order and imperfectly encodes equations; consequential formulas and figures were checked against rendered pages. Local PDF, readable text, and provenance and hashes.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.