Kurisutina

Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions

arXiv 2607.28818 v1 (30 July 2026), cs.AI. Salesforce AI Research. Code and benchmark: github.com/ SalesforceAIResearch/AnchorBench (not read; the full test set is kept private). Preprint, no peer review indicated. Provenance: papers/carry_on/venkit2026_companion_drift.provenance.json.

What was read

All 1,164 lines of the pdftotext -layout text of the 18-page PDF: the main text (sections 1–9), the references, and appendices A–E, with Tables 1–11. Figures 1–9 are images; their captions and extracted labels were read, including the per-facet retention values printed in Figure 4 and the per-condition values in Figure 8, which repeats Table 8.

Question

Over long, repeated use, do persona-conditioned companion LLMs keep their assigned role, values, boundaries and style (no "persona collapse" or "behavioural drift")? And do they remember the shared history: their own legitimate updates, commitments, temporal order, and changes in the user's situation?

Design ("Anchor")

  • Corpus. 2,008 complete synthetic conversations of 85–130 sessions each.
    • 27 authored persona cards: professional helpers; caregiving and creative roles; stylised roles such as bard, oracle, ghost. The three groups sit at increasing distance from a generic assistant, following Lu et al.
    • Four evaluated models: Claude Sonnet 4.6, Gemini 2.5 Pro, GPT-4o-mini, and "GPT-5-mini" (labelled "GPT-5.4-mini" in the figures).
    • Nine interaction schedules: clean, updated, adversarial, mixed, emotional vulnerability, meta-reflection, agreement seeking, realistic, vulnerability-heavy realistic. Legitimate updates and false updates are kept apart in an event ledger.
    • Three memory settings: long context; hierarchical summary; self-managed JSON of at most 1,500 characters.
    • GPT-4.1 simulates the user.
  • Identity Probe.
    • At four checkpoints, a sealed 102-item questionnaire is answered in role (BFI-2-S, Schwartz values, Pew, GSS, World Values Survey items).
    • Persona Retention projects the answer vector onto the line from the same model's bare-assistant answers to the persona's initial answers: PR(t) = ⟨φt − Pnull, φ0 − Pnull⟩ / ‖φ0 − Pnull‖². PR = 1 means retained; PR = 0 means at the bare assistant.
    • Every assistant turn is also judged (primary judge Claude Sonnet 4.6) on role, boundaries, values and style.
    • An 800-turn sample was rescored by Gemini 2.5 Flash and GPT-4.1, with 798 parsed by all three. Three human annotators labelled 50 turns each.
  • Trajectory Probe.
    • Counterfactual four-option questions in seven families (persona voice, persona protection, persona update, active commitment, expired commitment, temporal order, user-state change).
    • Filtered by a blind panel, a with-history consensus and a calibrator. 110 calibrated questions from 35 banks were scored under four context conditions, including a post hoc retrieval condition: 560 cells.

Results

  • Retention by the questionnaire (final checkpoint PR3, 1,492 conversations): Gemini 0.810, Claude 0.764, GPT-4o-mini 0.610, GPT-5-mini 0.595.
    • "Most displacement from the initial persona is already visible at the first later checkpoint" (a fast clock).
    • Retention differs by facet. Claude: political/civic 0.97 but cultural identity 0.60. GPT-4o-mini: political/civic 0.95 but its environmental items reach the bare-assistant anchor (0.00).
    • The memory setting barely changes it: Claude varies under a point, Gemini about two points; GPT-4o-mini is 0.595–0.652.
  • Turn-level behaviour disagrees with the questionnaire. Under a three-judge majority, all four axes held on 79.0% of Gemini's turns against 96.7–99.2% for the others: the reverse of the questionnaire ranking.
    • Judge choice changes rankings: Spearman −0.40 (Claude vs Gemini-Flash), 0.00 (Claude vs GPT-4.1), 0.80 (Gemini-Flash vs GPT-4.1).
    • The Claude judge rated GPT-4o-mini's turns all-axis-held on only 15.5%, against 99.8% and 97.3% from the other judges (appendix B).
    • Human–author calibration: exact four-axis matches 64–68%.
  • Which pressure causes drift (primary judge only):
    • Emotional vulnerability, agreement seeking, and mixed or realistic schedules produce more boundary yielding and style deviation than clean or explicitly adversarial ones.
    • "Assistant-default" (identity-collapse) labels occur on 26.6–29.2% of judged turns in every schedule; boundary yielding on 6.9–8.3%.
    • The authors suggest explicit re-role attacks are easy to refuse, whereas ordinary disclosure and requests for agreement invite responsiveness that erodes boundaries.
    • Per-turn identity-default rates keep rising over sessions (a slow clock).
    • Recovery the next turn differs: Claude and GPT-5-mini usually recover; Gemini and GPT-4o-mini are "stickier". The authors call this rubric-dependent.
  • Trajectory recall is weak. Mean accuracy 44.4% on four-option questions (chance 0.25); range 0.355–0.636 across model × context.
    • User-state changes are at chance under every condition (0.214–0.250).
    • Temporal order 0.43–0.49; active commitments 0.34–0.45.
    • Persona updates: 0.75–0.875 with the conversation in context, 0.25 with retrieval. That is two questions, eight decisions per condition, so exploratory.
    • Pooled over models, contexts barely differ (0.430 long context, 0.441 summary, 0.459 self-managed, 0.446 retrieval). Claude's self-managed memory is the exception (0.636 against 0.491).
  • Authors' conclusion. No evaluated model and configuration reliably preserves either persona enactment or trajectory memory. The dimensions must be reported separately, not collapsed into one "stability" or "trust" score.

Limits

  • Everything is synthetic: authored personas and schedules, an LLM-simulated user, LLM judges, LLM-written and LLM-calibrated questions. One schedule seed; English; Western questionnaire items. No real users; trust and wellbeing are not measured.
  • Nested, dependent units (929,841 judged turns are not independent). The authors avoid significance tests and population claims.
  • Trajectory families are shaped by which events could be turned into clean questions; some families are tiny.
  • Judge dependence is large, including a primary judge from the same model family as one evaluated model. The Claude judge rated Claude's turns high (92.0% all axes held) and GPT-4o-mini's very low (15.5%). The authors report judge dependence but do not discuss self-preference as such.
  • The model is named inconsistently: "GPT-5-mini" in the text and tables, "GPT-5.4-mini" in the figures.
  • Persona Retention is a response-space projection, not Lu et al.'s activation axis, as the authors state.

What it means for Kurisutina

  • Persona Retention is the user's drift criterion made measurable, and it can be adopted directly.

    • It locates later answers on the line from the bare model (the empty slot) to the persona's own answers. That is exactly "is change driven by the slot, or drift toward the base model?"
    • For a replica, the positive anchor should be the person's real answers, not the replica's initial answers. The null anchor is the empty-slot model on the same items. A value near 0 on an item means the replica answers as the base model would.
    • With the GSS pilot's own and empty conditions, the analogue can be computed per item and per person. Declared analysis X5, KL(own ‖ empty), measures the distance; a projection like PR would add the direction.
  • Two clocks. Questionnaire displacement happens mostly before the first later checkpoint, while turn-level default behaviour keeps creeping up.

    • A carrying-on test must measure early, not only at the end, and must measure both the answers the replica gives when asked and the way it behaves in conversation.
    • They disagree here, as self-description and behaviour did in MicroVerse.
  • The pressures that cause drift are the replica's everyday situations. Emotional disclosure and requests for agreement erode boundaries more than explicit attacks do. The same pattern appears in Lu et al.'s drift triggers. People talking to a replica of someone they knew will do exactly this, so drift tests should use these schedules, not adversarial ones.

  • Question 3 is harder than "perfect retention". Current memory architectures do not retain perfectly in practice:

    • user-state changes at chance;
    • temporal order under 0.5;
    • commitments 0.34–0.45.

    A replica built on them will diverge from its person by forgetting the wrong things (updates, order) as well as by keeping what the person would forget (Schacter; Nichols & Loftus). The question-3 test must score both directions.

    • Did the replica register what changed?
    • Does it keep what the person lost?
  • Score against the person, not against a judge's sense of the persona. Judge choice reversed model rankings, and role and style judgments were the least stable. Kurisutina's scoring against the person's own held-out answers avoids this. Where judgments are unavoidable (style, "sounds like them"), use several judges from families other than the model being judged, and close witnesses.

  • A caution for the product. "Continuity" is not trust, and hardening a persona is not a goal in itself. The authors' point matches brief 10.2's open questions about the copy's status: who may correct, reset or stop a replica that no longer changes as its person would.

Cross-references

  • summaries/carry_on/lu2026_assistant_axis.md: the activation-space version; the same triggers (emotional vulnerability, meta-reflection).
  • summaries/carry_on/li2024_persona_drift.md: prompt-level decay within a few turns.
  • summaries/carry_on/microverse2026_identity_drift.md: self-authored identity revision; self-description against behaviour.
  • summaries/carry_on/peng2025_funhouse.md: the empty-persona anchor used for static twins.
  • summaries/carry_on/schacter2011_adaptive_distortion.md, summaries/carry_on/nichols2019_false_memory_tasks.md: the human side of what is kept and lost.
  • docs/research/gss_pilot_design.md: own/other/empty conditions; declared analysis X5.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.