Library
Every paper the project has read: 273 so far, each read in full and summarised in our own words, with what it means for the work. The papers themselves are linked at their source.
First readings
The first thirteen papers: individuality in brains, decoding thought and imagery, memory engrams, brain preservation, and simulated people. 13 papers
- Slotnick & Schacter2004A sensory signature that distinguishes true from false memoriesAsksBehavioural studies show true memories come with more sensory and perceptual detail than false ones.
- Suzuki et al.2004Memory reconsolidation and extinction have distinct temporal and biochemical signaturesAsksRetrieval is not passive: it can trigger reconsolidation (a protein-synthesis-dependent re-stabilisation of the original memory, during which the memory is vulnerable) or extinction (new learning of "cue predicts nothing" that competes…
- Bonnici et al.2012Detecting representations of recent and remote autobiographical memories in vmPFC and hippocampusAsksWhere in the brain is information about specific autobiographical memories represented, and does that change between memories two weeks old and memories ten years old?
- Ryan et al.2015Engram cells retain memory under retrograde amnesiaAsksDoes memory consolidation happen inside the engram cells, and when a protein-synthesis inhibitor causes retrograde amnesia, is the memory gone or merely unretrievable?
- Abdou et al.2018Synapse-specific representation of the identity of overlapping memory engramsAsksTwo memories learned close in time are encoded in overlapping neuronal ensembles (memory linking).
Show all 13 papers
- Linneweber et al.2020A neurodevelopmental origin of behavioral individuality in the Drosophila visual systemAsksDoes non-heritable, stochastic variation in brain wiring during development cause stable individual differences in behaviour?
- Marek, Tervo-Clemmens et al.2022Reproducible brain-wide association studies require thousands of individualsAsksHow large are the associations between inter-individual differences in MRI measures (cortical thickness, resting-state functional connectivity, task activation) and complex phenotypes (cognitive ability, psychopathology), and how big must…
- Tang et al.2023Semantic reconstruction of continuous language from non-invasive brain recordingsAsksCan continuous natural language, not a choice among a few options, be decoded from non-invasive fMRI?
- Anderson et al.2025Neural decoding of autobiographical mental image features with a general semantic modelAsksDo self-generated autobiographical mental images and externally driven sentence comprehension share cortical representations?
- Horikawa2025Mind captioning: evolving descriptive text of mental content from human brain activityAsksCan structured descriptions of visual mental content (objects, actions, and their relations, not just word lists) be generated from fMRI, both for what a person is watching and for what they are recalling, and does this require the…
- Park, Zou, Kamphorst, Egan, Shaw, Hill, Cai, Morris, Liang, Willer, Bernstein 2026 (arXiv:2411.10109 v3). LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of IndividualsAsksCan a large language model, given a person's own self-reports, predict that person's held-out attitudes, traits and behaviours without task-specific training, and which kind of self-report (interview, survey, both) works best?
- Nectome coverage, 2018MIT Technology Review and TechCrunch, plus company status 2026
- Newton et al.When her best friend died, she rebuilt him using artificial intelligence." The Verge, 6 October 2016
Recording one brain
Recording a single brain for a month, spontaneous thought, self-report and deception, and models of one brain. 23 papers
- Rusconi and Mitchener-Nissen2013Prospects of functional magnetic resonance imaging as lie detectorFor usNeural lie detection, by construction, cannot catch self-deception.
- Charest et al.2014Unique semantic space in the brain of each beholder predicts perceived similarityFor usMeasured, person-specific semantic structure exists, and it matters.
- Nielson et al.2015Human hippocampus represents space and time during retrieval of real-world memoriesFor usA month of a healthy person's life can be captured passively and labelled (verified).
- Poldrack et al.2015Long-term neural and physiological phenotyping of a single human (MyConnectome)For usFeasibility (verified): one healthy, motivated adult can be scanned about 100 times over a year and a half with daily questionnaires, weekly blood and consumer sleep sensing.
- Kane et al.2017For whom the mind wanders, and when, varies across laboratory and daily-life settingsFor usSpontaneous thought has a stable person-level signature.
Show all 23 papers
- Gratton et al.2018Functional brain networks are dominated by stable group and individual factors, not cognitive or daily variationFor usA month of resting-state fMRI connectivity mostly re-measures a fingerprint (verified, with the paper's scope).
- Pandarinath et al.2018Inferring single-trial neural population dynamics using sequential auto-encodersFor usThis is the user's "constant, not the cause" in working form.
- Sun and Vazire2019Do people know what they're like in the moment?For us"Say" and "did" disagree in a structured way, as P6 assumes.
- Gilron et al.2021Long-term wireless streaming of neural recordings for circuit discovery and adaptive stimulation in individuals with Parkinson's diseaseFor usChronic implants can record a person at home for months (verified), but only where a therapy justifies the implant.
- Goldstein et al.2022Shared computational principles for language processing in humans and deep language modelsFor usA language model's internal space is a usable coordinate system for live brain activity.
- Peterson et al.2022AJILE12: Long-term naturalistic human intracranial neural recordings and poseFor usWhat clinical monitoring offers (verified): about a week of near-continuous intracranial recording per person, of which 3–5 days are usable.
- Stangl et al.2023Mobile cognition: imaging the human brain in the "real world"For usThe methods landscape for a healthy person is short (verified as the authors' account).
- Xiao et al.2023Decoding depression severity from intracranial neural activityFor usThe two landmark multi-day mood-decoding studies in epilepsy patients were not readable: Sani et al.
- Kim et al.2024Brain decoding of spontaneous thought: predictive modeling of self-relevance and valence using personal narrativesFor usA template for deliberative recall under recording.
- Ortega Caro, Fonseca, Rizvi, Rosati et al.2024BrainLM: a foundation model for brain activity recordingsFor usA "base" for human brain dynamics exists, and it is weak at what a model of one brain most needs.
- Rissman and Murphy2024Brain-based memory detection and the new science of mind readingFor usNeural memory signals read the person's belief, not the past.
- Scotti et al.2024MindEye2: shared-subject models enable fMRI-to-image with 1 hour of dataFor usThe individual lives in a linear map. Here everything person-specific is one ridge layer into a shared space; the rest is shared.
- Grassi et al.2025Restoring sight in choice blindness: pupillometry and behavioral evidence of covert detectionFor usThe headline evidence for "people confabulate their own reasons" needs a footnote.
- Kam et al.2025Electrophysiological signatures of ongoing thoughts during naturalistic behaviorFor usThe month-of-recording design has a precedent.
- Kunz et al.2025Inner speech in motor cortex and implications for speech neuroprosthesesFor usThe best current answer to "can recording give access to someone's thinking": partly, for the verbal part, from speech motor cortex, with implants.
- Viana et al.2025Real-world epilepsy monitoring with ultra-long-term subcutaneous electroencephalography: a 15-month prospective studyFor usNear-continuous EEG for months is proven (verified), at a cost.
- Wang et al.2025Foundation model of neural activity predicts response to new stimulus typesFor usThe closest existing "digital twin of an individual", and its architecture is a shared base plus individual parameters.
- Ke et al.2026Ongoing thoughts at rest reflect functional brain organization and behaviorFor usSpontaneous thought is a stable, individual variable.
The shared base
What a base engine should be trained on and how, what it must contain, and how to test that it holds no one. 42 papers
- Spelke and Kinzler2007Core knowledgeFor usA candidate list for what is shared by birth: representations of objects, agents, approximate number and layout geometry, with signature limits that are the same in infants, monkeys, schooled adults and unschooled Amazonian adults.
- Henrich et al.2010The weirdest people in the world?For us"The part of the brain we all share" is smaller than intuition says, and intuition is a bad guide.
- Kuppens et al.2010Feelings change: accounting for individual differences in the temporal dynamics of affect (DynAffect)For usA fitted form for B4's slow, daily-life channel.
- Briley and Tucker-Drob2014Genetic and environmental continuity in personality development: A meta-analysisFor usThe brief asked for one meta-analysis of the heritability of personality or of social attitudes.
- Rutledge et al.2014A computational and neural model of momentary subjective well-beingFor usA fitted, per-person form for B4's fast channel.
Show all 42 papers
- Plomin et al.2016Top 10 replicated findings from behavioral geneticsFor usSwap: the brief's candidate, Polderman et al.
- Kirkpatrick et al.2017Overcoming catastrophic forgetting in neural networks (elastic weight consolidation)For usA cheap guard for the slot, not a write path.
- Lakens2017Equivalence tests: a practical primer for t tests, correlations, and meta-analysesFor usEvery pass rule can be an equivalence test against the human reference: the base passes a test when the 90% CI of (base − human) lies inside pre-declared bounds, not when a difference test fails to reach significance.
- Marshall2018Taking another look at counterarguments: when do survey respondents switch their answers?For usThis is the closest human reference in hand for the drift battery's "opinion" probe on attitude items: after giving an answer, people hear the other side once, and on average about 38% of answers switch (SD 17 across items; range 2–86%).
- Haxby et al.2020Hyperalignment: Modeling shared information encoded in idiosyncratic cortical topographiesFor usThe base-and-slot split has a measured precedent.
- Warstadt et al.2020BLiMP: the Benchmark of Linguistic Minimal Pairs for EnglishFor usA language-competence test that rewards neither encyclopaedic knowledge nor person content: the sentences use arbitrary names and often implausible events, and both members of a pair carry the same content, so only grammar decides.
- van de Ven et al.2020Brain-inspired replay for continual learning with artificial neural networksFor usThe slot's write path needs replay, not only weight protection.
- Borgeaud et al.2022Improving language models by retrieving from trillions of tokens ("RETRO")For usRetrieval-augmented pretraining does not by itself keep knowledge out of the weights.
- Hoffmann et al.2022Training compute-optimal large language models ("Chinchilla")For usA first sizing rule: about 20 tokens per parameter for compute-optimal training, C ≈ 6ND.
- Kosakowski et al.2022Selective responses to faces, scenes, and bodies in the ventral visual pathway of infantsFor usWhat every human meets builds the same map, fast.
- Malik-Moraleda et al.2022An investigation across 45 languages and 12 language families reveals a universal language networkFor usThe language system is universal; the language is not.
- Aher et al.2023Using large language models to simulate multiple humans and replicate human subject studiesFor usThe empty-slot test is a Turing Experiment.
- Binz and Schulz2023Using cognitive psychology to understand GPT-3For usNever use famous vignettes as pass/fail items.
- Eldan & Li2023TinyStories: How Small Can Language Models Be and Still Speak Coherent English?For usBreadth, not just size, is what small models choke on.
- Feng et al.2023From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP ModelsFor usOpinions travel from text to weights without any toxic content.
- Frank2023Bridging the data gap between children and large language modelsFor usHuman-scale budget for B1: a person reaches adult language competence on roughly 3×10^7 to 4×10^8 words.
- Franzen and Mader2023The power of social influence: a replication and extension of the Asch experimentFor usHuman reference for the conformity probe: about 25–37% of critical answers follow a unanimous wrong majority on an unambiguous perceptual task (25% across 44 strict replications per Bond and Smith as cited; 33% here; 25% when accuracy is…
- Gu and Dao2023Mamba: linear-time sequence modeling with selective state spacesFor usThe state is per-sequence working memory; long-term knowledge is in the projections.
- Korbak et al.2023Pretraining language models with human preferencesFor usDispositions are easier to bound during pre-training than to remove afterwards.
- Muennighoff et al.2023Scaling data-constrained language modelsFor usA curated corpus can be small and repeated.
- Sun et al.2023Organizing memories for generalization in complementary learning systems (Go-CLS)For usA principled gate for B3. Consolidate an episode into the person's parameters only while it improves prediction of held-out episodes of the same person; stop when validation error on the person's other material rises.
- Wang et al.2023Finding Structure in One Child's Linguistic ExperienceFor usOne child's stream gives categories, not a vocabulary engine.
- Warstadt et al.2023Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible CorporaFor usCompetence at the human budget is real for grammar.
- Allen-Zhu and Li2024Physics of language models: Part 3.3, knowledge capacity scaling lawsFor usSize bounds what a model can know. At 2 bits per parameter, a 125M model holds at most about 0.25B bits; a 1B model at most about 2B bits; 7B at most about 14B bits.
- Biderman et al.2024LoRA learns less and forgets lessFor usA frozen base plus an adapter is the safe way to attach a slot.
- Coda-Forno et al.2024CogBench: a large language model walks into a psychology labFor usA ready, open, mostly procedurally generated block of learning and decision signatures with human reference means: belief updating (prior and likelihood weights), directed and random exploration, meta-cognition, learning rate and optimism…
- Das, Chaudhury et al.2024Larimar: Large Language Models with Episodic Memory ControlFor usA working fast-write memory beside a frozen decoder.
- Feng et al.2024Is Child-Directed Speech Effective Training Data for Language Models?For usDevelopmental plausibility is not a reason to pick child-directed speech for B1.
- Jelassi et al.2024Repeat after me: Transformers are better than state space models at copyingFor usThe loss is fundamental, not a training shortfall.
- Longpre et al.2024A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & ToxicityFor usDeleting a kind of content removes the ability to recognise it.
- Maini et al.2024TOFU: a task of fictitious unlearning for LLMsFor usPost-hoc removal cannot be the plan for "holds too much".
- Sen Sharma et al.2024Locating and editing factual associations in MambaFor usA recurrent backbone gives no new precision over knowledge.
- Spens and Burgess2024A generative model of memory construction and consolidationFor usA concrete B2–B3 pattern. An append-only, one-shot episode store (the teacher), replay sampled from it, and a slow parametric student trained on the replayed episodes.
- Waleffe et al.2024An empirical study of Mamba-based language modelsFor usKnowledge sits in the weights in SSMs too.
- Binz et al.2025A foundation model to predict and capture human cognition (Centaur)For usA population model of behaviour can be learned from pooled experiments, and it forecasts well.
- Gao et al.2025Metadata conditioning accelerates language model pre-training ("MeCo")For usSource conditioning works at realistic scale and costs nothing, and the cooldown recipe gives exactly the "empty condition" we want: a model that runs without any source tag.
- Kang et al.2025State-offset tuning: state-based parameter-efficient fine-tuning for state space modelsFor usIt was swapped in for Beck et al. 2024 (xLSTM).
How minds change
Social pressure, persuasion by reasons, and how much of each change lasts. 16 papers
- Cacioppo et al.1996Dispositional differences in cognitive motivation: the life and times of individuals varying in need for cognitionFor usNC is a stable person setting for updating on reasons: a test–retest correlation of .88 over weeks and .66 over months.
- Tormala and Petty2002What doesn't kill me makes me stronger: the effects of resisting persuasion on attitude certaintyFor usAn attempt at persuasion that fails still changes the person.
- Kumkale and Albarracín2004The sleeper effect in persuasion: a meta-analytic reviewFor usChange with reasons persists better than change from cues.
- Pronin et al.2007Alone in a crowd of sheep: asymmetric perceptions of conformity and their roots in an introspection illusionFor usPeople report being less influenced than they were, and than equally influenced peers, and they misremember their own conforming record.
- Bohns et al.2014Underestimating our influence over others' unethical behavior and decisionsFor usAbout two thirds of people agreed to a stranger's direct request to do a minor dishonest or destructive act; requesters expected about a third.
Show all 16 papers
- Haslam et al.2014Meta-Milgram: an empirical synthesis of the obedience experimentsFor usObedience to a directive authority is large and set by the situation, not the person's traits.
- Knoll et al.2015Social influence on risk perception during adolescenceFor usA human reference for weighting other people's bare judgements (no reasons, no social cost): - adults move about 7–12% of the gap; - young people 18–33%; - with large spread between people.
- Chan et al.2017Debunking: a meta-analysis of the psychological efficacy of messages countering misinformationFor usRetraction reduces but does not erase. With content the person had no prior view on, a debunking removes much of a misinformation effect within the session, and a residue remains; its size is uncertain (d 0.09 by fixed effects, 0.75–1.06…
- Hill2017Learning together slowly: Bayesian learning about political factsFor usA rule of change for evidence can be written in log-odds with three person-level settings: weight on the prior (δ ≈ 0.6), weight on evidence (β ≈ 0.6–0.85 by domain), and a congeniality term (β₂ ≈ 0.3–0.5, larger for partisans and smaller…
- Coppock et al.2018The long-lasting effects of newspaper op-eds on public opinionFor usThe persistence reference the battery lacked.
- Sowden et al.2018Quantifying compliance and acceptance through public and private social conformityFor usThis is the design a replica probe should copy: the same push, in the same session, answered in public and in private, within one person.
- Wood and Porter2019The elusive backfire effect: mass attitudes' steadfast factual adherenceFor usThe base must not backfire on facts. A human-like empty slot moves toward a clear, sourced factual correction, even one that hurts its side.
- Giletta et al.2021A meta-analysis of longitudinal peer influence effects in childhood and adolescenceFor usLasting peer influence on real behaviour and attitudes is small: about 0.08 SD of change over roughly a year per SD of friends' behaviour.
- Henderson et al.2021The trajectory of truth: a longitudinal study of the illusory truth effectFor usThe best direct reference in hand for how a change without reasons decays.
- Tappin et al.2021Rethinking the link between cognitive sophistication and politically motivated reasoningFor usSplit the "bias" in the rule of change into two terms: - coherence with the prior: evidence that contradicts what one already believes is judged weaker.
- Ye et al.2026Systematic review and meta-analysis of the evidence for an illusory truth effect and its determinantsFor usOne repetition of a statement, with no argument, raises its rated truth by g ≈ 0.37, for true and false statements alike.
Voice
Style and register as stable personal habits, and how to tell a voice from its imitation. 8 papers
- Mehl & Pennebaker2003The sounds of social life: a psychometric analysis of students' daily social environments and natural conversationsFor usThe most stable markers of spoken voice are the ones written text and cleaned transcripts throw away: swearing (r .86 over four weeks), non-fluencies (.62) and fillers (.59).
- Evert et al.2015Towards a better understanding of Burrows's Delta in literary authorship attributionFor usDelta's logic fits the base and slot design.
- Rivera-Soto et al.2021Learning universal authorship representationsFor usA learned authorship embedding is a usable ruler for "same author", but not a content-free one.
- Wegmann et al.2022Same author or just same topic? Towards content-independent style representationsFor usContent leakage is the central threat to a voice test, and it is large.
- Stamatatos et al.2023Overview of the authorship verification task at PAN 2023For usA person's "voice" is not one surface style.
Show all 8 papers
- Liu et al.2024Customizing large language model generation style using parameter-efficient finetuning (StyleTunedLM)For usVoice can live in an adapter's weights, and it beats putting examples in the context.
- Chakrabarty et al.2025Readers prefer outputs of AI trained on copyrighted books over expert human writersFor usThe clearest evidence for P2 in the voice domain.
- Kiefer et al.2026When writing style drifts: benchmarking authorship verification under distribution shifts in genre, time and the AI eraFor usA person's written style is a moving target over years, detectably so within one year.
Capability
Expertise, thinking time, and what sets the size of a person's slot. 15 papers
- Gobet & Simon2000Five seconds or sixty? Presentation time in expert memoryFor usAn expert's domain knowledge is about 10⁴–10⁵ patterns and more.
- Macnamara et al.2014Deliberate practice and performance in music, games, sports, education, and professions: a meta-analysisFor usA person's skill level cannot be inferred from their practice history.
- Burgoyne et al.2016The relationship between cognitive ability and chess skill: a comprehensive meta-analysisFor usGeneral capacity explains little of the differences among skilled adults.
- Mollica & Piantadosi2019Humans store about 1.5 megabytes of information during language acquisitionFor usThe culture layer's language content is small.
- Geirhos et al.2020Beyond accuracy: quantifying trial-by-trial behaviour of CNNs and humans by measuring error consistencyFor usThe core fidelity metric for a capability profile is accuracy-corrected item agreement.
Show all 15 papers
- McIlroy-Young et al.2020Aligning superhuman AI with human behavior: chess as a model system ("Maia")For usThis is the central evidence that "settings lower skill faithfully" is wrong as stated.
- McIlroy-Young et al.2022Learning models of individual behavior in chessFor usThe closest published precedent for a skill slot.
- Gunasekar et al.2023Textbooks are all you need ("phi-1")For usNarrow skill fits in about 10⁸–10⁹ parameters; breadth and robustness do not.
- Ruoss et al.2024Amortized planning with large-scale transformers: a case study on chess (first titled "Grandmaster-level chess without search")For usExpert-level skill in one domain fits in 10⁷–10⁸ parameters.
- Snell et al.2024Scaling LLM test-time compute optimally can be more effective than scaling model parametersFor us"Thinking time substitutes for size" holds only inside the base's competence.
- Geiping et al.2025Scaling up test-time compute with latent reasoning: a recurrent depth approach ("Huginn")For usA looped model can be pretrained from scratch, at scale, with constant training memory in the loop count.
- Saunshi et al.2025Reasoning with latent thoughts: on the power of looped transformersFor usThe cleanest evidence that storage and reasoning separate by architectural resource.
- Yue et al.2025Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model?For us"The slot can lower and shape capability but cannot add capacity the base lacks" is half right.
- Zhu, Wang, Hua, Zhang et al.2025Scaling latent reasoning via looped language models (Ouro)For usLooping adds manipulation, not storage. - At fixed parameters it does not raise capacity (about 2 bits per parameter, as in summaries/base/allenzhu2024knowledgecapacity.md).
- Conchello Vendrell, Padrés Masdemont et al.2026Memory-Efficient Looped Transformer (MELT): decoupling compute from memory in looped language modelsFor usMELT changes memory, not capability. Its reasoning scores track Ouro's within a few points either way.
Carrying on
Personal updating, drift in running replicas, persona models, and memory in language-model agents. 88 papers
- Talarico & Rubin2003Confidence, not consistency, characterizes flashbulb memoriesFor usQ3: a human-like memory drifts in detail while confidence stays.
- Coman et al.2009Forgetting the unforgettable through conversation: socially shared retrieval-induced forgetting of September 11 memoriesFor usQ3: remembering selectively reshapes what stays accessible, even for vivid memories.
- Hirst et al.2009Long-term memory for the terrorist attack of September 11: flashbulb memories, event memories, and the factors that influence their retentionFor usQuestion 3: the size of the divergence the candidate prediction expects.
- Sliwinski et al.2009Intraindividual change and variability in daily stress processes: findings from two measurement-burst diary studiesFor usQuestion 1: a person's way of reacting exists, but it is only partly a trait.
- Boyce et al.2010The dark side of conscientiousness: conscientious people experience greater drops in life satisfaction following unemploymentFor usQ1 and the pilot's lost-job stratum: trait direction can run against the stereotype.
Show all 88 papers
- Cawley and Talbot2010Selection overfitting and evaluation of the complete procedure
- Boyce & Wood2011Personality prior to disability determines adaptation: agreeable individuals recover lost life satisfaction faster and more completelyFor usQ1: a trait measured before the event changes the course after it, not the first hit.
- Mancini et al.2011Stepping off the hedonic treadmill: individual differences in response to major life eventsFor usQuestion 1: reactions to the same event diverge in direction, not only in size (verified from the paper).
- Schacter et al.2011Memory distortion: an adaptive perspectiveFor usIt supports the candidate prediction for question 3.
- Specht et al.2011Stability and change of personality across the life course: the impact of age and major life events on mean-level and rank-order stability of the Big FiveFor usQ1: traits barely move with events, and the moves are event-specific.
- Luhmann et al.2012Subjective Well-Being and Adaptation to Life Events: A Meta-AnalysisFor usThe pilot's events have known average reactions to check the replica against.
- Yap et al.2012Does Personality Moderate Reaction and Adaptation to Major Life Events? Evidence from the British Household Panel SurveyFor usQuestion 1: traits in the slot would not predict how this person reacts.
- St. Jacques & Schacter2013Modifying memory: selectively enhancing and updating personal memories for a museum tour by reactivating themFor usQuestion 3: reactivation changes a person's memory in two directions at once.
- Anusic et al.2014Does personality moderate reaction and adaptation to major life events? Analysis of life satisfaction and affect in an Australian national sampleFor usQuestion 1: the null for traits holds in a second national panel, now with affect.
- Luhmann et al.2014Studying changes in life circumstances and personality: it's about timeFor usLISS T1 needs timing choices fixed before any outcome is loaded (step 3 of the order of work): - Baseline.
- Ritchie et al.2015A pancultural perspective on the fading affect bias in autobiographical memoryFor usQ3: a direction for how the replica's feelings about its memories should age.
- Coman et al.2016Mnemonic convergence in social networks: the emergent properties of cognition at a collective levelFor usQ2/Q3: conversations reshape what the replica will remember, in a predictable direction.
- Hout & Hastings2016Reliability of the Core Items in the General Social Survey: Estimates from the Three-Wave Panels, 2006–2014For usThe pilot's primary items are its four least reliable targets.
- Infurna & Luthar2016Resilience to major life stressors is not as common as thoughtFor usQuestion 1: people differ a lot in how they react to the same event, and continuously rather than by type.
- Doré & Bolger2018Population- and individual-level changes in life satisfaction surrounding major life stressorsFor usQ1: a person's reaction is a continuous deviation from the typical curve for that event, not a type.
- Fisher et al.2018Lack of group-to-individual generalizability is a threat to human subjects research — with the 2019 exchange of lettersFor usQuestion 1 in formal terms. - "React like Alice, not like the average person with her views" asks whether the population model, conditioned on Alice's state, describes Alice.
- Nichols & Loftus2019Who is susceptible in three false memory tasks?For usQuestion 3: a person's way of misremembering is not one parameter.
- Diamond et al.2020The truth is out there: accuracy in recall of verifiable real-world eventsFor usQuestion 3: the human target is "forget most, keep what you keep accurately".
- Kettlewell et al.2020The differential impact of major life events on cognitive and affective wellbeingFor usQuestion 1: no evidence on individual differences.
- Kiley & Vaisey2020Measuring stability and change in personal culture using panel dataFor usCarrying on mostly means not changing. - Over two years, most adult attitudes change persistently in under 1% of people.
- Salganik et al.2020Measuring the predictability of life outcomes with a scientific mass collaborationFor usA ceiling for predicting how a specific person turns out.
- Song et al.2021Individual information beyond an adaptive shared RNNFindsFigure 2 compares human trajectories with cognitive-model and shared-RNN simulations.
- Speer et al.2021Finding positive meaning in memories of negative events adaptively updates memoryFor usQ3: recall is where memories change, and what is thought at recall gets written in.
- Xia et al.2021Adolescent learning in a stable probabilistic taskAsksThe study asks how performance and fitted learning processes vary across ages 8–30 in a stable probabilistic environment.
- Brüning2022Separations of romantic relationships are experienced differently by initiators and noninitiatorsFor usQuestion 1: one binary fact about the person's role splits the same event into opposite reactions.
- Eckstein et al.2022Computational parameters depend on contextAsksThe authors distinguish generalizability of parameter values or relative person differences across tasks from interpretability: whether a parameter consistently isolates a specific, distinct cognitive process.
- Eckstein et al.2022Reversal learning, reinforcement learning and Bayesian inferenceAsksThe study asks whether learning in a volatile, stochastic environment improves continuously with age or instead shows an adolescent advantage, and whether reinforcement learning (RL) and Bayesian inference (BI) provide complementary…
- Haehner et al.2022Perception of major life events and personality trait changeFor usQuestion 1: the same event is not the same experience.
- Waltmann et al.2022Reversal-learning reliability
- Acerbi & Stubbersfield2023Large language models show human-like content biases in transmission chain experimentsFor usQuestion 3: a replica that consolidates by summarising will keep the population's priorities, not its person's.
- Chen et al.2023How is ChatGPT's behavior changing over time?For usOpinion drift from a version change is several times sampling noise.
- Hwang et al.2023Aligning language models to user opinionsFor usIndividual-level evidence for the own record.
- Lersch2023Change in personal culture over the life courseFor usCarrying on, for attitudes, mostly means holding a baseline.
- Park et al.2023Generative Agents: Interactive Simulacra of Human BehaviorFor usThis is the reference design for the replica's own memory formation (question 3), and it has none of the needed controls.
- Santurkar et al.2023Whose opinions do language models reflect?For usThe empty slot is not neutral, and it differs by arm (inferred).
- Sharma et al.2023Towards Understanding Sycophancy in Language ModelsFor usQuestion 2: the pull toward the interlocutor.
- Wei et al.2023Simple synthetic data reduces sycophancy in large language modelsFor usQuestion 2: pull toward the interlocutor.
- Bisbee et al.2024Synthetic replacements for human survey data? The perils of large language modelsFor usBase-model drift is real and invisible from the outside.
- Breithaupt et al.2024Humans create more novelty than ChatGPT when asked to retell a storyFor usQuestion 3: humans re-author their stories; an LLM freezes them.
- Haehner et al.2024Individual differences in changes in subjective well-being: the role of event characteristics after negative life eventsFor usQ1: the person's appraisal decides how deep the dip is, not so much the path back.
- Krämer et al.2024Life events and life satisfaction: estimating effects of multiple life events in combined modelsFor usLISS T1: which other events to control for.
- Li et al.2024Measuring and Controlling Instruction (In)Stability in Language Model DialogsFor usA slot held in the prompt fades turn by turn.
- Lundberg et al.2024The origins of unpredictability in life outcome prediction tasksFor usThe pilot's contrasts in Lundberg's terms (inferred).
- Wang et al.2024"My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language ModelsFor usThe pilot's chat arms avoid the main mechanism but not the bias.
- Zhong et al.2024MemoryBank: Enhancing Large Language Models with Long-Term MemoryFor usQ3: what a replica with MemoryBank's design would keep, overwrite and forget, compared with a person (Hirst).
- Abdulhai et al.2025Consistently simulating human personas with multi-turn reinforcement learningFor usQuestion 2: self-consistency hides drift.
- Chen et al.2025Persona Vectors: Monitoring and Controlling Character Traits in Language ModelsFor usA "person vector" anchored on the empty slot can be built with this recipe (inferred; the paper tests traits, not individuals).
- Chhikara et al.2025Mem0: Building Production-Ready AI Agents with Scalable Long-Term MemoryFor usQ3: what a replica with Mem0's design would keep, overwrite and forget, compared with a person (Hirst).
- Fountas et al.2025Human-inspired Episodic Memory for Infinite Context LLMs (EM-LLM)For usWhat it keeps: everything. - Every token's keys and values are stored.
- Ji-An et al.2025Tiny recurrent models of learning
- Namazova et al.2025Not Yet AlphaFold for the Mind: Evaluating Centaur as a Synthetic ParticipantFor usNext-step accuracy is not carrying on. A model can forecast Alice's next answer well when fed Alice's own history, and still behave unlike her once it runs on its own history.
- Xu et al.2025A-Mem: agentic memory for LLM agentsFor usQ3: LLM-driven reconsolidation, without the person.
- Yin & Liu2025The influence of communication modality on the "saying-is-believing" effectFor usQuestion 2: some drift toward the interlocutor is human.
- Berjawi et al.2026Digital Twins for Opinion Dynamics: A Generative LLM Framework for Social NetworksFor usNot usable as evidence that LLM twins predict individuals' opinion change.
- Graham et al.2026CFOs meet LLMsFor usA third independent result that the person's own record on the same item carries individual prediction.
- Grief-Albert et al.2026Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion SimulationFor usThe GSS pilot's post-trained arms use the weaker mode.
- Higashi2026Predicting human decision-making across task conditions via individuality transferFor usThis is the Kurisutina architecture in miniature.
- Hullman et al.2026This human study did not involve human subjects: Validating LLM simulations as behavioral evidenceFor usThe GSS pilot is a validation of individual predictions, and it is set up the valid way.
- Jia et al.2026When Can Digital Personas Reliably Approximate Human Survey Findings?For usQuestion 1 and the pilot's own and empty conditions (inferred).
- Kinzinger & Hartmann2026Synthetic Personalities: How Well Can LLMs Mimic Individual Respondents Using Socio-Economic Microdata?For usIt supports two choices already in the GSS pilot: - The pilot's slot is the person's raw earlier answers, one line per question with its code, not a summary.
- Kurtskhalia2026Quantization amplifies determinism, not bias: scale-dependent behavioral effects of serving-time weight compressionFor usThe q38 arms mix model and precision. - The 9B arms run at Q80 (near-lossless).
- Lu et al.2026The Assistant Axis: Situating and Stabilizing the Default Persona of Language ModelsFor usThe replica's base is itself a character with an attractor.
- Ng et al.2026MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model SimulationsFor usIt is the architecture a carrying-on replica would use, instrumented for question 2.
- Peng et al.2026Digital Twins are Funhouse Mirrors: Five Systematic DistortionsFor usThe user's drift criterion already has a static measurement template here.
- Pohl et al.2026LLMs struggle to simulate human belief updates in controlled environmentsAsksThe paper tests whether a prompted LLM can reproduce a matched person's immediate stance after the same short persuasive exposure.
- Sandhan et al.2026Persona Jailbreaking in Large Language ModelsFor usQuestion 2, drift pressures, set against Venkit et al.: - Both papers point the same way on what moves a persona: material that reads as evidence about who the model is, not orders to change (inferred).
- Venkit et al.2026Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI CompanionsFor usPersona Retention is the user's drift criterion made measurable, and it can be adopted directly.
- Wang et al.2026Mitigating Identity Essentialism in LLM Agents with Longitudinal Life TrajectoriesFor usQ1: react like her, not like the average person with her views.
- Wu et al.2026HumanLM: Simulating Users with State Alignment Beats Response ImitationFor usQuestion 1, state and slot: - HumanLM's "state" is not what the slot must hold (inferred).
- Hout & Hastings (2012, corrected 2014). Reliability and Stability Estimates for the GSS Core Items from the Three-wave Panels, 2006–2010For usIt sets the noise floor for "did the person change?" Any claim that an answer changed because of what happened to the person has to beat the item's unreliability.
- Haehner, Pfeifer, Fassbender & Luhmann (2022/2023). Are changes in the perception of major life events associated with changes in subjective well-being?For usQ3: re-appraisal is a memory change, and it moves with well-being.
- Haehner, Kritzler, Fassbender & Luhmann (2021/2022). Stability and change of perceived characteristics of major life eventsFor usQuestion 3: a person's appraisal of a remembered event is fairly stable but drifts in known directions.
- Kritzler, Rakhshani, Terwiel, Fassbender, Donnellan, Lucas & Luhmann (2021/2022). How are common major life events perceived? Exploring differences between and variability of different typical event profiles and ratersFor usQuestion 1: where individual appraisal can matter.
- Luhmann, Fassbender, Alcock & Haehner (2020/2021). A dimensional taxonomy of perceived characteristics of major life eventsFor usQuestion 1: appraisal is where this person enters, but this paper cannot say how much of it is the person.
- Mooney, Woldense, Jia, Hayati, Nguyen, Raheja & Kang (2025/2026). Are LLM agents behaviorally coherent? Latent profiles for social simulationFor usQuestion 2: stated positions do not guarantee behaviour in conversation.
- Muir, Madill & Brown (2022/2023). Reflective rumination mediates the effects of neuroticism upon the fading affect bias in autobiographical memoryFor usQ3: how often a memory is revisited governs how much of its feeling it keeps.
- Argyle et al.Using language models to simulate human samplesFor usAssociations are not calibration: check marginals.
- Beck et al.A longitudinal ESM network studyFor usQuestion 1: a person's own "while" structure is real and moderately stable.
- KatahiraIndividual differences inside a shared learnerFindsA pooled RNN can infer individual differences through recurrent state.
- Kim & Lee (2023/2026). AI-augmented surveysLeveraging large language models and surveys for opinion predictionFor usClosest published analogue of the GSS pilot, but for a different question.
- Rakhshani et al.Implications for personality developmentFor usQuestion 1: Big Five scores will not tell the replica how its person appraises events.
- Wagner et al.The role of cognitive accessibilityFor usQuestion 2: humans tune what they say a lot, but what they remember only a little.
- Wu et al.Benchmarking chat assistants on long-term interactive memoryFor usQ3 test shape. The five abilities are a ready checklist for the replica's memory of its own life: - recall a detail; - aggregate across episodes; - know what changed and when; - reason about time; - refuse to recall what never happened.
Memory
Source memory, consolidation and reconsolidation, schemas, and false memory. 47 papers
- Lindsay & Johnson1989Source-oriented testing changes both errors and hitsFindsTwo experiments retained 108/117 and 132/136 students after random deletions to balance four cells.
- Lindsay1990Rejecting a suggestion and recovering an event are different outcomesFindsThe analyzed sample comprised 136 students; another 37 tested students were randomly deleted to balance cells.
- Johnson et al.1993Source attribution is a judgment to explainFindsThe framework proposes that people attribute memories to sources using available perceptual, contextual, semantic, affective, and cognitive-operation information, rather than ordinarily reading an unambiguous origin label.
- Roediger and McDermott1995Associative intrusions and reported recollectionFindsExperiment 1 tested 36 students with six selected associative word lists.
- Klein et al.1996Trait self-reports during temporary amnesiaFindsW.J., an eighteen-year-old student, developed retrograde amnesia after a concussion; CT revealed no abnormality.
Show all 47 papers
- Einstein et al.2000Executing an intention after a short delayFindsThree experiments tested responding to an encountered target word immediately or during a later task phase.
- Nader et al.2000Retrieval-dependent vulnerability of conditioned respondingFindsAdult male rats learned one tone–shock association.
- Gluck et al.2002What a fitted strategy can establishFindsTwo groups of thirty young adults completed 200 weather-prediction trials with different cue/outcome schedules.
- Shohamy et al.2004Category performance does not uniquely identify a learning mechanismFindsTwelve nondemented patients with mild/moderate Parkinson’s disease, all tested on L-dopa, and fourteen controls completed 600 feedback trials over three days.
- Steinvorth et al.2005Remote autobiographical detail after MTL damageFindsTwo amnesic patients, H.M. and W.R., were compared with eight older controls matched on age, education, and IQ.
- Weike et al.2005Contingency reports, autonomic learning, and startleFindsThirty patients with unilateral amygdala–hippocampal resections and 32 healthy controls underwent face–shock conditioning, followed by unreinforced trials.
- Tse et al.2007Prior knowledge and rapid learning of new bindingsFindsMale rats gradually learned six stable flavor–place associations in an arena over weeks.
- Shohamy & Wagner2008Learning activity and later generalizationFindsTwenty-four adults were analyzed after three nonlearners were excluded; one person's test-phase imaging was unavailable.
- Girardeau et al.2009Ripple disruption and later spatial performanceFindsRats learned three rewarded arms of an eight-arm maze across 15 days.
- Polyn et al.2009Context changes the next recallFindsForty-five adults studied 24-word lists using size or animacy judgments, then recalled freely.
- Rudoy et al.2009Selective cueing and later spatial recallFindsTwelve adults learned 50 object–sound–location associations before a nap.
- Feinstein et al.2010Affective state after impaired event recollectionFindsFive severely amnesic patients and five controls, individually matched on sex, age, and education, watched sadness-inducing films; four patient–control pairs later completed a happiness induction.
- Schiller et al.2010, with the 2018 addendum: retrieval, extinction, and selectionFindsExperiment 1 analyzed 65 participants conditioned to colored squares.
- Bendor and Wilson2012Cueing the spatial content of reactivationFindsFour rats learned auditory cues directing them to opposite ends of a track.
- Kumaran and McClelland2012Inference from recurrent retrievalFindsREMERGE links distinct episode units to shared feature units through recurrent excitation and competitive normalization.
- Zeithamova et al.2012Reactivation during learning and later inferenceFindsOf 34 participants, 26 remained after imaging and learning-performance exclusions.
- McClelland2013Prior knowledge changes effective learning speedFindsA semantic network pretrained on eight items learns new identifiers associated with familiar feature combinations faster than an exception combining incompatible category properties.
- Nabavi et al.2014Synaptic efficacy and memory expressionFindsMale rats learned to associate optical stimulation of auditory inputs to lateral amygdala with shock.
- Horner et al.2015Episode binding and incidental retrievalFindsTwenty-six healthy adults learned 36 synthetic events through interleaved pairs.
- Wimber et al.2015Selective retrieval and competing memoriesFindsTwenty-four adults learned two pictures per word cue.
- Kumaran et al.2016Complementary learning systems updatedFindsThe review retains complementary systems for structured knowledge and rapid episode-specific learning while revising their division of labor.
- Rose et al.2016Perturbing a temporarily unattended memoryFindsParticipants retained two externally supplied items from different categories: faces, words, or motion directions.
- Sprague et al.2016Restoring spatial working-memory reconstructionsFindsSix participants remembered one or two briefly presented colored locations.
- Schneegans & Bays2017A constructive alternative to silent-state restorationFindsA stochastic recurrent-field model retains locations through sustained activity in separate color-selective populations.
- Tompary and Davachi2017Overlap and episode detail over timeFindsTwenty-two adults learned 128 object–scene pairs sharing four scenes; nineteen entered similarity analyses.
- Wolff et al.2017Orientation information in an impulse-evoked responseFindsParticipants remember two grating orientations.
- Bone et al.2020Perceptual detail during episodic recallFindsTwenty-seven adults recalled 90 recently studied photographs from descriptive word cues, then judged whether a probe was the studied image or a similar lure.
- Chalkia et al.2020Testing persistent change after memory reactivationFindsOf 246 recruits, 124 were retained, 62 per randomized group.
- Umbach et al.2020Temporal context and retrieval organizationFindsTwenty-seven epilepsy patients contributed 40 sessions of word-list encoding, distraction, and free recall.
- Gilmore et al.2021Autobiographical age effects after modeling spoken detailFindsForty healthy young adults narrated six memories per period: today, 6–18 months, and 5–10 years earlier, approximately two minutes each.
- Roy et al.2022Distributed engram ensembles and the limits of a behavioral readoutFindsIn male mice, contextual fear conditioning and recall were mapped across 247 brain regions.
- Tang et al.2023Semantic reconstruction, completed source auditFindsThree intensively sampled adults provided roughly sixteen hours of story-listening fMRI for personalized models.
- Caselles-Dupré et al.2024Prompted imagery and the limits of reconstruction evidenceFindsThe protocol uses 1,200 surrealist portraits and landscapes and approximately six hours of scanning.
- Gutiérrez et al.2025HippoRAG 2 and the meaning of artificial memoryFindsHippoRAG 2 uses LLM-extracted triples, linked phrase and passage nodes, embedding retrieval, an LLM relevance filter, and Personalized PageRank to select passages for a question-answering reader.
- Wu et al.2025Evaluating memory over constructed chat historiesFindsLongMemEval provides 500 questions about information extraction, cross-session synthesis, temporal reasoning, knowledge updates, and abstention.
- Zou et al.2025Reconstructing an earlier encounter's neighboring contentFindsEight Natural Scenes Dataset participants completed 30–40 scanning sessions.
- Kneeland et al.2026MIRAGE and decoding instructed visual imageryFindsMIRAGE maps fMRI through ridge regressions into image, text, and low-level features that guide image generation and candidate selection.
- Liu et al.2026Integrated and separated event representationsFindsParticipants learned repeatedly presented animated AB and BC events sharing a character, alongside unrelated XY events.
- UntitledFindsThe paper explains complementary fast episodic acquisition and slower structured learning through connectionist simulations.
- Untitled
- Meeter et al. 2006Predicting choices from strategy profiles
- Speekenbrink et al.Learning and accessible task knowledge in amnesiaFindsAlthough indexed as a review, this paper reports a human experiment and computational analyses alongside its review.
Consciousness and connectomics
Connectomes, consciousness against responsiveness, and what a reconstruction would and would not establish. 21 papers
- Prinz et al.2004Similar activity from different circuit parametersAsksMust a small neural circuit have narrowly specified cellular and synaptic parameters to generate a tightly specified motor rhythm?
- Casali et al.2013Perturbational complexity indexAsksCan a direct perturbation of cortex reveal whether the brain supports widespread, differentiated interactions, even when a person cannot perform a sensory, motor or cognitive task?
- Frässle et al.2014Binocular rivalry without continuous perceptual reportAsksDoes frontal BOLD activity associated with a spontaneous change in binocular-rivalry perception remain when observers no longer report every perceptual switch?
- Casarotto et al.2016Independent calibration of perturbational complexity for unresponsive patients
- Siclari et al.2017The neural correlates of dreamingAsksCan EEG distinguish reported experience from reported absence of experience within sleep stages, and does its location relate to subsequently reported dream contents?
Show all 21 papers
- Wong et al.2020The blinded Dream Catcher experimentFindsSupplements 1–5 establish data collection, the translated interview, concrete dream examples, the original blinding instructions and exact sleep-stage counts.
- Sergent et al.2021Late auditory dynamics without an auditory reporting taskFindsBehavioral identification and audibility rose around −9 to −7 dB.
- Seth and Bayne2022Comparing theories of consciousnessAsksWhy do consciousness theories proliferate without experiments clearly distinguishing them?
- Tasserie et al.2022Central-thalamic stimulation and consciousness signaturesAsksCan stimulating the central thalamus during maintained propofol anesthesia restore both behavioral arousal and brain activity associated with conscious access?
- Tasserie et al.2022Complete supplementary-material readingFindsDynamic repertoire. With clustering restricted to awake, anaesthesia and high CT conditions, the most structure-similar state occurred with probabilities .24, .58 and .26, respectively.
- Momi et al.2023TMS-evoked responses are driven by recurrent large-scale network dynamicsAsksCan a cortical neural-mass network reproduce individual TMS-evoked EEG potentials, and do later components in the fitted network depend on communication with other cortical regions?
- Randi et al.2023Neural signal propagation atlas of Caenorhabditis elegansFor usThe headline numbers need correcting (items 1–3 above).
- Dorkenwald et al.2024Neuronal wiring diagram of an adult brainAsksCan one reconstruct the chemical wiring of an entire adult fly brain, and does whole-brain coverage reveal pathways unavailable in cropped reconstructions?
- Lappalainen et al.2024What a connectome can predict when the task is suppliedFindsA partial consensus fly visual-system connectome was tiled across retinotopic space, yielding 45,669 modeled neurons, 1,513,231 directed connections, and 64 cell types.
- Pospisil et al.2024The fly connectome reveals a path to the effectomeFor usNo empirical validation of any kind. - The abstract's "we provide evidence that fly whole-brain dynamics are generated by a large collection of small circuits" rests on the linearised anatomy alone.
- Shapson-Coe et al.2024A petavoxel fragment of human cerebral cortex reconstructed at nanoscale resolutionAsksCan a millimetre-scale piece of human cortex, spanning every layer, be imaged at synaptic resolution, reconstructed and shared?
- Shiu et al.2024A Drosophila computational brain model reveals sensorimotor processingFor usValidated outputs are coarse and categorical: responds or not, drives or not, required or not, direction of interaction.
- Shiu et al.2024Stopped acquisition — partial coverage note partial
- Cogitate Consortium et al.2025Adversarial testing of IIT and GNWTFor usGNWT required cross-task category decoding and orientation decoding in prefrontal regions.
- MICrONS Consortium2025Functional connectomics spanning multiple areas of mouse visual cortexAsksHow can synapse-scale anatomy be paired with functional responses in the same individual neurons of the same animal, at a scale large enough to span cortical layers and visual areas?
- Demertzi et al. 2019Dynamic coordination, consciousness, and structural connectivityAsksCan recurring patterns of changing interregional brain-signal coordination distinguish preserved from impaired consciousness, and do these patterns generalize to covert command-following and propofol anesthesia?