arXiv 2608.19621, version 3 (22 September 2026). v1 (20 August) and v2 (21 August 2026) were not read.
Tsinghua University and Quan Cheng Laboratory. A preprint with no venue stated, under the arXiv non-exclusive
licence. The code (github.com/halsayxi/LifeMem) was not read. Provenance:
papers/carry_on/wang2026_lifemem.provenance.json.
This paper came out of the search on the strength of its abstract. On reading, it does not meet the search
criterion. The criterion was prediction of later answers or of change, not same-wave prediction; see Method and
Limits. For that reason summaries/carry_on/jia2026_liss_personas.md was read as well.
What was read
- All 1,383 lines of
pdftotext -layoutoutput (20 pages): main text, limitations, appendices and Tables 1–11. The appendices cover the WVS check, dataset construction, prompts, baselines, metric definitions, Algorithm 1 and additional results. - The pages are two-column, and the layout text interleaves the columns line by line. It was read with that in mind.
- Figures 1–8 are images. Only their captions and the labels pdftotext recovered were read; the plotted values were not viewed. Figures 3 and 5 (retrieval depth, event coverage) hold their values only in the plots, and their trends are given here as the text states them.
- The code was not read.
Question
LLM agents given a demographic profile are more alike within a group, and more different between groups, than real people. The authors call this "identity essentialism": the model treats "group-average tendencies as individual traits".
Does it help to give each agent its own longitudinal life history, in a retrieval memory and in a per-agent LoRA adapter? The test is whether simulated populations then match humans in distribution, diversity and wave-to-wave transitions.
Method
-
Motivating check (WVS wave 7). 2,000 respondents; Llama-3.1-8B-Instruct was given each demographic profile. The silhouette of the three SES groups in answer space was −0.02 for humans and 0.19 for the agents.
-
Data. Understanding Society (UKHLS) public-use files, waves 1–15 (2009–2024).
- 14,104 people appear in all 15 waves.
- 100 were sampled (seed 42); seeds 43 and 44 were used for robustness.
-
Variables. Rule-based semantic matching sorts items into three kinds (Table 5, per wave: 61–71 demographic, 37–89 event and 11–64 evaluation variables).
- Demographic: they build the static profile.
- "life events": ordinary survey question–answer pairs, for example whether the household can afford friends round for a drink or meal, or whether a child attends a playgroup. gpt-3.5-turbo rewrote each as a second-person and a first-person statement.
- Evaluation: for example, would you prefer to move, and do you feel closer to one party. A separate common longitudinal set of 50 variables present in all waves (housing, employment, income, caregiving, family, education) is used for the transition metric.
-
LifeMem.
- Structured memory. Events are appended and never deleted. Retrieval scores cosine similarity (all-MiniLM-L6-v2) times exp(−0.105 × waves ago) and takes the top 5.
- Parametric memory. One LoRA adapter per person, rank 8, on the q, v, o and down projections; the base model is
frozen.
- Each wave the adapter is trained on that wave's events, both question→answer and statement reconstruction.
- Up to 4 older events are replayed at weight 0.5.
- An L2 penalty pulls the adapter toward its previous state (γ = 0.01).
-
Protocol (verified from the implementation section and Algorithm 1).
- At each wave the adapter is first updated on the events observed at that wave, and the events are added to memory.
- Then the same wave's evaluation questions are answered.
- So the prediction for wave t uses the person's wave-t answers to other questions. No later wave is forecast.
-
Baselines.
- Direct: the question only.
- Profile: demographics.
- Multilingual and Anti-Stereotype prompting.
- SimVBG: a generated backstory.
- Full History: all events through the current wave, most recent first, within 4,096 tokens.
- Event RAG: bge-m3 retrieval, top 5.
- Random Event: 5 events taken from other agents.
-
Models. Llama-3.1-8B-Instruct, Ministral-3-8B-Instruct-2512 and Qwen3.5-9B (thinking off). Temperature 0, one answer per agent per question, bf16.
-
Metrics, all population-level.
- KL(human ‖ model) per wave–question cell, with ε = 10⁻⁹ added to the counts;
- the gap in within-group pairwise difference, by demographic group;
- the gap in normalised entropy;
- transition-distribution Jensen–Shannon divergence between the pooled (answer at t, answer at t+1) pairs of humans and of agents, on the common set.
No metric compares an agent's answer with its own person's answer. Significance comes from paired t-tests over questions.
Results
-
Table 1 (seed 42; values rounded from the table's four decimals). LifeMem is best on every metric for every model.
Metric Llama Ministral Qwen3.5-9B KL: LifeMem 3.45 1.97 2.69 KL: Profile 7.15 6.83 5.32 KL: Direct 15.37 15.49 16.00 Within-group gap: LifeMem 0.289 0.228 0.252 Within-group gap: Profile 0.384 0.362 0.346 Transition JS: LifeMem 0.333 0.340 0.306 Transition JS: Profile 0.389 0.361 0.347 - For Ministral, LifeMem's transition JS is not significantly better than Profile, Anti-Stereotype, SimVBG, Full History or Random Event.
-
Ablations (Table 2). Removing either memory worsens most metrics. For Ministral's transition JS, neither ablation is significantly worse.
-
Own events against other people's events (Table 1; the differences are my computation). Random Event minus Event RAG, for Llama, Ministral and Qwen:
Metric Llama Ministral Qwen KL +0.27 +0.39 +0.23 within-group gap +0.027 +0.027 +0.036 entropy gap +0.055 +0.064 +0.064 transition JS +0.016 −0.077 +0.007 - Own events beat other people's modestly on the static metrics, and not on the change metric.
- There is no LifeMem arm trained on other people's events.
-
Robustness (Table 8, seeds 42–44). LifeMem stays best on KL, within-group gap and entropy gap. Transition JS is not reported across seeds.
-
Retrieval depth (Fig. 3, Llama). Gains flatten beyond K = 20 while runtime and tokens keep rising.
-
LoRA states. In a PCA plot (Fig. 4) they spread apart over the waves.
-
Cost.
- An adapter update takes 2.63 s per agent and wave (Llama).
- Inference takes 154.5 ms per question plus 1.4 ms to load the adapter: 10.7 times Direct.
-
Validity. Humans give a valid answer in 64.31% of cells and the models in 89–93%. Only cells where both are valid are scored.
Limits
- Not a forecast.
- Wave-t answers are generated after training on the person's wave-t answers to other items.
- The "within-person response change across life stages" in the abstract is a population statistic over wave-to-wave pairs. It was produced with same-wave information.
- Population metrics cannot tell whose history drives which agent (verified from the metric definitions).
- KL, entropy and transition JS pool over respondents. Permuting which person's history feeds which agent leaves them unchanged, except for which cells are retained.
- Only the within-group gap uses the agent's own demographic group.
- So the paper cannot show that an agent represents its person. That would need a control that trains LifeMem on shuffled trajectories, and there is none.
- No population baseline (inferred).
- Sampling answers from the human marginals, or from the human transition table, would score near zero on KL and transition JS without modelling anyone. No such baseline is reported.
- There is no no-change baseline either.
- The size of the KL is set by ε (my computation).
- With greedy decoding and one identical prompt, Direct gives every agent the same answer.
- A label the model never uses gets P ≈ 10⁻¹¹. Each 0.1 of human mass on such a label then adds about 2.3 to the KL (natural log).
- Direct's 15–16 therefore mainly says that a large share of human answers, roughly half or more, fall on labels the model never produced. The paper does not state the log base.
- Differences in KL between methods mostly measure how many answer labels each method covers.
- Scale and scope.
- 100 people, deterministic decoding, one answer per agent, and only instruction-tuned 8–9B models.
- Most "life events" are routine items, not events.
- Only exact targets are excluded from the inputs; near-duplicate items between inputs and targets are not screened.
- Inconsistencies found:
- The implementation section calls bge-m3 "the stronger" retriever for Event RAG. In Table 11 (Ministral), all-MiniLM-L6-v2 beats BGE-M3 for Event RAG on all three metrics (KL 3.845 against 4.334).
- Table 3's relative latencies are not all time ÷ 14.4 ms (my computation). The table gives SimVBG 13.7, Event RAG 66.5 and Random Event 5.6; the division gives 13.6, 66.4 and 5.5.
- The common set has 50 variables "available in every wave", but Table 5 lists only 11 evaluation variables in waves 1 and 2. So the common set is counted separately from Table 5, which the text does not say.
- The abstract and conclusion claim better "within-person response change". The only change metric is not significant for one of the three models, and it is absent from the robustness table.
- Checked:
- Table 2's "LifeMem without Parametric Memory" for Ministral (3.8454) equals Event RAG with MiniLM in Table 11 (3.845);
- λ = 0.105 per wave equals the Event RAG decay of 0.9 per wave (exp(−0.105) = 0.900);
- Table 4's range of wave sizes is 41,601–77,495.
What it means for Kurisutina
- Q1: react like her, not like the average person with her views.
- The essentialism diagnosis is our failure mode, measured: demographic conditioning compresses people toward their group.
- With greedy decoding and no persona, the model gives everyone one answer (Direct). That is the empty slot at its
extreme.
- The pilot's empty condition reads probabilities, so it avoids this collapse.
- Its chat arms can still be very sharp: q38 was overconfident in the smoke run.
- The proposed remedy is not shown to individuate. The metrics would not notice agents built from the wrong person (verified above).
- The pilot's own / other / empty conditions (inferred).
- Random Event is this paper's "other" condition, but only for prompt retrieval and only in distribution.
- Own events gained little over others' events (KL 0.23–0.39; transition JS mixed).
- The pilot's D2 is the per-person version, and it is the right instrument. Parity between own and other in distribution says nothing about D2.
- A table baseline is missing here (inferred). The pilot has P(later | earlier) and P(later | earlier, event). This paper would need the human transition table to show its transition JS beats copying the population. Any population-level metric Kurisutina reports for replicas should include the table as a baseline.
- A measure to adopt: transition-distribution JS.
- It is a cheap population-level drift check for replicas run over several simulated waves. Compare the replica population's pooled t→t+1 table with the humans'.
- It complements the phi test from Kiley & Vaisey.
- Read it beside per-person scores, never instead of them.
- A condition to adopt (proposal).
- A per-person LoRA slot on a frozen base is a concrete architecture for the goal's rule that change must come from the slot, not from base drift. The base never changes; only the person's adapter does, with an explicit penalty on how far it moves per update.
- The proper test, which this paper did not run:
- train on waves ≤ t and forecast wave t+1;
- compare the own adapter, another person's adapter and no adapter;
- score each person against the no-change and table baselines.
- Kim & Lee's trained respondent embedding is the cross-sectional precedent.
- Q3: memory formation (inferred).
- This is a complementary-learning-systems design. The store retains everything: events are appended and never deleted, and decay affects only retrieval priority. The adapter overwrites, but under regularisation.
- Forgetting as loss is not modelled. So such a replica would keep what the person may have lost, which is the direction of divergence seen in Hirst et al.
- Model overlap. Qwen3.5-9B with thinking off is one of the three backbones, the same family as the pilot's chat9b arm.
Cross-references
summaries/carry_on/jia2026_liss_personas.md: later answers of the same people (LISS), scored per person.summaries/carry_on/kim2024_ai_augmented_surveys.md: a trained per-person slot on a frozen backbone.summaries/carry_on/kiley2020_personal_culture.md: persistent and reverting change; phi as a drift test.summaries/carry_on/santurkar2023_opinionqa.md,summaries/carry_on/bisbee2024_synthetic_replacements.md: collapse to modes; caricature.summaries/carry_on/zhong2023_memorybank.md,summaries/carry_on/chhikara2025_mem0.md: other memory designs.summaries/carry_on/hirst2009_911_memory.md: what people lose and keep.docs/research/gss_pilot_design.md: D2 and the table baselines.