Research
The architecture, the first mechanism test, six research batches and the studies that came before them. Every number on this page has a record in the repository, and negative results are reported as results.
Goal
A replica of one specific person that carries on as them: built from what their living memory can retrieve, recognisably them, and continuing to learn and change as they would after it is switched on.
We count functional equivalence as being the person. Once activated, the replica branches from the original. It does not need to track the original's later life; it needs to think and act like them, and to change the way they would. How far the replica sits from the original is a measurement we make, not the goal itself.
Architecture
The base supplies capability and never identity. The person lives in a slot of weights and state, and new experience reaches it only through gated, offline consolidation (Fig. 1). Seven principles constrain every component:
| Principle | What it rules out | |
|---|---|---|
| P1 | The base supplies capability, not identity. | Opinions, a persona, or knowledge the person does not have. |
| P2 | The person lives in weights and state, not in text. | A replica that becomes a blank assistant when its store is emptied. |
| P3 | Function, not mechanism. | Claims about wiring that no measurement can check. |
| P4 | Memory is a gated, living system. | A conversation rewriting the person directly. |
| P5 | Judge on learning equivalence. | Agreement before an experience without agreement on the change after it. |
| P6 | Triangulate words, actions and brain data. | Treating self-report as ground truth; self-deception is stored as "believed by the person". |
| P7 | The base is defined by a frozen test battery. | Calling something a base because of its architecture. |
M1, the synthetic world
Real people's beliefs cannot be read off directly, so the first test runs in a world where they can. M1 generates persons whose beliefs, and the way those beliefs change, are known exactly, and asks whether a person held in a learned slot on an engine we train ourselves can read their own memories, stay themselves, and carry on.
| Part | What it is |
|---|---|
| World | 8 regions, 64 towns, 32 guilds and 1,536 residents, with two inference rules (a resident's region and workplace follow from their town and guild; spouses share a town) and 256 shared culture facts. |
| Persons | 512 for training and 48 for evaluation, each with values, susceptibility to pressure, credulity, ties and a voice; about 80k tokens of first-person life each. |
| Reference mind | An exact symbolic mind scores every answer. A learner fitted to the person's own past sets the ceiling, and a population average sets the floor. |
| Engine | A 36.8M-parameter decoder with a cross-attention reader over up to 20 retrieved memory units, trained on 400M tokens. |
| Gate | Reading, person conditioning, and a small lifecycle, each with declared bars. A cell without data counts as missing, never as passed, and the stop rule at 100M and 200M tokens can only stop a run. |
The first corpus trained without trouble and then failed its first real test: 216,787 of 407,140 reads failed (agreement with the reference mind below 0.5), because training rarely needed the memory store. The facts had simply been memorised. The corpus was rebuilt so that reading is necessary, and every revision since has closed a way for the gate to pass without the capability it names. Two independent reviewers check each revision. The engine is pretraining again now.
Revision historyResearch batches
Each batch reads its papers in full and ends in a synthesis with decisions.
| Batch | Question | Decision |
|---|---|---|
| R1 | Can one brain be recorded for a month? | Not continuously. A dense-sampled month instead; good labels matter more than recording hours. |
| R2 | Can brain data separate true from false memories? | No. Memory signals track belief, not history, so triangulation rests on records. |
| R3 | Can a model of one brain learn its dynamics? | The individual is a small adapter on a shared model. Every slot part must beat another person's version and a cheaper channel. |
| R4 | What should the base be trained on? | Three layers (base, culture, slot). Knowledge is controlled with files and gates, because unlearning is unreliable. |
| Spec | What must the base be? | People share processes, not content. No off-the-shelf language model qualifies as the base. |
| R5 | How do people change under influence, and what is voice? | Pressure moves what people say far more than what lasts. Voice lives in weights. |
| R6 | What sets capability, and the sizes? | Expertise is content: roughly 10⁷–10⁸ slot parameters per domain, learned from the person's own decisions. |
Studies
Before building our own engine we tested the open questions on existing models and public data.
| Study | Data | Result |
|---|---|---|
| Drift battery | A 9B chat model prompted as 37 survey respondents; 1,510 rows | Within a session the probability of an interlocutor's stated answer rose by 0.83; memory notes carried 79–90% of that into the next session (Fig. 2). |
| GSS panel pilot | Three-wave General Social Survey panels; 5 model arms, 36,810 forecasts | Every model arm forecast later answers worse than a population table that knows the life event: Brier score 0.577–0.615 against 0.548 (lower is better). |
| Personal learning, context | 267 people, 32,154 scored choices | Personal summaries made a shared model's predictions worse: log loss up 0.0375 nats per choice; 77 of 267 people improved. |
| Personal learning, retest | 40 people, two reversal-learning sessions | A personal response policy helped (log loss down 0.039 nats per choice; 31 of 40 people improved); history added to an adaptive model did not. |
| Belief updating (Pohl 2026) | 391 people, argument responses | No detectable gain from earlier personal reactions. |
| Learning transfer (Higashi 2026) | Fresh-cohort replication, 576 fits | A stable gain over both strong comparators is not established. |
| Neural increment | Reproduction of Anderson et al. (2025) | Added to the person's other ratings and their own description, a brain recording cut the mean squared error of predicting held-out ratings (0–6 scale) by 0.08%, from 2.204 to 2.202; p = 0.35. |
| Slot in weights (E2) | Smoke tests, two persons | The adapter resisted pressure but also ignored real evidence: it moved 0.33 as far as the prompt-held teacher (+0.156 against +0.465), where 0.5 was required. |
Method
- 01
Read every line
A paper counts as read only when all of it has been read from the full text. Its written summary then stands as our record of it, with the source's hash kept beside it.
- 02
Label every claim
Verified (run or traced, with a file reference), inferred, or guessed.
- 03
Never move a criterion
No test, threshold or acceptance clause is loosened to get a pass, including by a pre-declared "if it fails, demote it" rule. If a check looks wrong, work stops and a person decides.
- 04
Review independently
Two reviewers read each design and each build without seeing each other's review. A defect that could let a check pass for the wrong reason blocks the run.