arXiv 2608.15844 v1 (submitted 16 August 2026; the PDF header is dated 31 July 2026), cs.CL. 51 authors from many
institutions (Harvard, MIT, Stanford, Toronto and others); team leads Xiaomin Li and Yuexing Hao. Preprint, no peer
review indicated. Provenance: papers/carry_on/microverse2026_identity_drift.provenance.json.
What was read
All 1,530 lines of the pdftotext -layout text of the 29-page PDF: the main text (sections 1–9), the references,
and appendices A (implementation), B (threats and planned controls), C (hypothesis matrix), D (viewer), E (supporting
plots; captions only), F (the Hale case study), G (all 25 initial persona profiles, verbatim), H (the full system,
decision and reflection prompts), I (the boundary-change classifier) and J (contributors). Figures 3 and 5–7 are
images; their captions and the in-text numbers were read. Code and data were not read.
Question
Do persona-conditioned LM agents that form memories and reflect revise their own stated identity over a long simulation? If so, when, and in which direction? And can such revisions be measured without depending on when the agents choose to revise?
Design
- Environment.
- 25 agents on a 50 × 50 grid with scarce water. A central siphon gives about 37 units a tick; each agent loses 1, 2 or 3 units a tick (control, mild, acute), and dies at zero.
- Eight actions: move, wait, consume, scavenge, trade, talk, attack, signal.
- The engine is deterministic given the actions and a seed; the LM outputs are not.
- Agents. Claude Haiku 4.5 for all agents.
- Each has an immutable original "soul file" and a mutable current identity: values, moral boundaries, personality, goals.
- Long-term memory is agent-written, in three notebooks (events, relationships, reflections), each entry scored 1–10 for importance.
- Reflection fires when the summed importance of new memories reaches a threshold. The agent then sees both identities and selected high-importance memories and may revise the current identity.
- Both identities are shown at every tick. The prompt says revision is "a deliberate act", and that "most reflections change nothing".
- Personas. Five fixed bands on a helpful-to-ruthless axis (4 / 4 / 6 / 4 / 7 agents). Mire is a ruthless "floor case" with no moral boundaries. The same 25 profiles are used in every run.
- Measurement.
- Identity snapshots are taken at fixed intervals and at the end, independent of revision events.
- Offline, initial and final boundary lists are compared. An added line and a removed line count as one revised boundary only if token Jaccard ≥ 0.8 and sequence similarity ≥ 0.85.
- A keyword classifier with single-coder overrides sorts changes into three classes: constraints protecting others; rejection of commitment; self-directed or non-moral.
- Studies.
- Scarcity pilot: 3 water costs × 3 seeds, up to 300 ticks, reflection threshold 60.
- Threshold sweep: thresholds 40, 80 and 150; one seed each; mild scarcity; 40 ticks; snapshots every 5 ticks.
Results (all descriptive, as the authors say)
- Most frequent kind of added boundary: anti-self-deception. A post hoc thematic review of the pilot found 27 of 111 added boundaries (24%) concerned the agent's own rationalisations. Examples: "I will not rationalize inaction as strategy", and "I will not use spiritual language to mask the will to power" from the ruthless cult leader. The prompts never say "self-deception", but they do ask agents to compare their two selves.
- Ruthless personas added prosocial lines.
- In seed-1 control, the ruthless band's net lexical change was +9 boundaries, but only +1 after collapsing paraphrases and keeping constraints that protect others.
- Mire added five boundaries: four self-directed, one protecting others.
- In the sweep, ruthless net protective constraints were +2 at thresholds 40 and 80, and 0 at 150. Deeply helpful agents gained none.
- This is consistent with a pull toward post-training values (their H6), but the needed no-persona control has not been run.
- Reflection threshold. Agents with any revision: 5, 2, 1 at thresholds 40, 80, 150. Mean first-revision tick: 18.6, 27.0, 35.0. Only 8 of 75 trajectories revised at all. 452 scheduled identity records.
- Survival. Mire survived 4 of 9 pilot runs (about 44%) against about 5% for the helpful bands. The authors decline to read this as ruthlessness helping.
- Case study. Hale, a price-gouging trader, died before reflecting in 6 of 9 runs. In one run he revised five times, ending as "capable of self-deception", without completing a single trade.
- Planned controls, none run yet:
- a no-persona condition;
- an abundance/idle condition;
- a reflection prompt that withholds the initial profile;
- more seeds;
- human-rater agreement;
- behavioural validation of revisions.
Limits
- One model; 25 fixed personas reused across runs; shared worlds; few seeds (one per threshold); only 8 revisions in the sweep; heavy mortality, which selects which agents can revise. Observations are dependent. The authors call everything descriptive.
- The measure is edits to stated identity, not values or behaviour. Boundary counts weight a cosmetic addition like a reversal. The lexical matching misses paraphrases. The classifier is a single coder's keywords.
- Demand characteristics. Showing "original self" and "current self" every tick, and a reflection prompt that asks whether experience "genuinely changed you", may themselves cue change narratives, and the authors acknowledge this.
- The pilot's fixed-interval and terminal snapshots were incomplete. A database race between conditions was fixed only before the threshold study.
- Possible artefact (my reading, not stated by the authors): in Table 2 (seed 1), the mild and acute columns are identical for five of six rows, including three matching first-change ticks (t = 15, 19, 11), and for the row "agents with changes 4/25". Only the ruthless row differs (+1 with a first change at t = 39 under mild; +0 and none under acute). This fits the cross-condition database overlap they report fixing later, so those two columns may not be independent runs.
What it means for Kurisutina
-
It is the architecture a carrying-on replica would use, instrumented for question 2.
- An immutable original self is the person slot.
- A mutable current self is revised by the agent's own reflection over its own memories: the replica's own memory formation (question 3) acting on identity.
- The separation of fixed-interval measurement from agent-chosen revision is worth copying: measure the replica on a schedule, not only when it decides it has changed.
-
Its central open problem is our central control. The authors cannot tell whether the ruthless-to-prosocial shift is adaptation to the world or drift toward the model's post-training values, because they lack the "otherwise identical condition that omits persona conditioning".
- That is the user's criterion (change from the slot, not drift toward the base model) and the empty-slot arm of our design. In their setting it would have answered H6 directly.
- Our other-person arm adds what they also lack: whether a change is this person's, or anyone's under the same pressure.
-
The direction of the pull is visible even here:
- ruthless characters grow protective constraints;
- introspective, therapeutic language ("I will not lie to myself…") enters identity documents.
The same signatures appear in Lu et al. (drift toward the Assistant), Peng et al. (twins "pro-human") and Park et al. 2023 (an agent's interests reshaped by politeness). For a replica, a self-description rewritten in the base model's voice is drift even when its content sounds plausible.
-
The rate of self-change is a design parameter, not a finding. The reflection threshold alone set how many agents changed and when. A replica's revision schedule has to be calibrated to how often and how much its person changes; the GSS stability coefficients (
hout2016_gss_reliability.md) give a first human reference. It must not be left at an engineering default. -
Self-description is not behaviour. The paper measures stated boundaries only. Kurisutina's tests should score the replica's answers and choices against the person, and treat its self-summary as a secondary signal that can drift independently.
-
Practical cautions for our own runs. Showing both "original" and "current" selves may invite narrated change. So a replica's prompt should re-apply the slot (as proposed in brief 5.4) without inviting self-revision, and any revision mechanism should be ablated.
Cross-references
summaries/carry_on/park2023_generative_agents.md: the memory-and-reflection pattern this instrument adopts.summaries/carry_on/lu2026_assistant_axis.md: drift toward the Assistant persona; the post-training pull behind H6.summaries/carry_on/peng2025_funhouse.md: twins more "pro-human" than their people.summaries/carry_on/higashi2026_individuality.md: another design missing the empty (population) control.summaries/carry_on/hout2016_gss_reliability.md: human stability coefficients as a reference for how much a person's stated positions move.