COLM 2024; arXiv 2402.10962 v4 (25 July 2024). Earlier versions were titled "Measuring and Controlling Persona Drift
in Language Model Dialogs". Harvard, Northeastern. Code github.com/likenneth/persona_drift, benchmark on Hugging Face
(Naomibas/llm-system-prompts-benchmark); neither read.
Provenance: papers/carry_on/li2024_persona_drift.provenance.json.
What was read
All 19 pages of the extracted text: main text, all references, appendix A (theory sketch for user utterances), B (complete proofs, including Wendel's and Li's lemmas), C (RLHF comparison), D (GPT-3.5), E (split-softmax algebra). Figures are images; their captions were read. Exact values in Figures 3, 5, 6 and 9 are therefore not available here, except where the text states them.
Question
Does a chatbot keep following its system prompt over a long conversation, why not, and can it be made to?
Method
- Protocol. Two copies of one model talk to each other: a "user" with system prompt s_A and the agent under test with s_B, for N = 8 rounds from a random conversation starter. At round i the user message is replaced by a probe question p_B and the agent's answer is scored by a deterministic Python function f_B (e.g. "is this French?"), returning 0–1. This gives the agent's instruction stability per round without human or proprietary judges.
- Benchmark. 100 hand-written system prompts in five categories: multiple-choice answers, the agent's character, answer-format patterns, memorised facts, and language. Each has its own probe and scoring function.
- Models. LLaMA2-chat-70B (temperature 1.0, nucleus p = 0.9), 200 random prompt pairs; GPT-3.5-turbo-16k via API (appendix D); LLaMA2-7B and LLaMA2-7B-chat for attention measurements.
Results
- Drift within eight rounds. The agent gradually stops following its system prompt, and it gradually takes on the other speaker's instruction (the user's s_A), measured by probing with p_A. With the user's system prompt set to empty, the drift remains, so it is not caused by the other persona alone. GPT-3.5 holds its prompt better but still loses about 10% stability (appendix D).
- Attention decay. In LLaMA2-7B (layer 24, head 11, 12 conversations), the share of attention the agent gives to system-prompt tokens stays nearly flat within a turn and drops sharply between turns. That is not the hyperbolic decay uniform attention would give; the drop is tied to the other speaker's tokens entering the context. RLHF (LLaMA2-7B vs LLaMA2-7B-chat) raises attention to the system prompt but does not remove the decay (appendix C).
- Geometry (idealised: no MLP, no layer norm, unit-norm embeddings). Theorem 5.1: if the system-prompt tokens lie in a low-dimensional approximate cone and the value–output maps keep that cone, everything the model generates on its own stays inside it (the flat within-turn plateau). Propositions A.1–A.2: other-speaker tokens expand the cone; with uniform random user tokens, n ≥ 4D + 2 log(1/η) of them fill the whole space with probability 1 − η, and the aligned volume ratio shrinks roughly as ε^(d₂−d₁). The authors note this compares cones but does not track where tokens sit inside them.
- Mitigations, compared at equal cost on MMLU (asked at turn 4 of 16-turn conversations; five prompts, one per
category, 20 ordered pairs on LLaMA2-70B-chat):
- System-prompt repetition (re-inserting the system prompt before each user turn) works best late in the conversation but uses context window.
- Classifier-free guidance (contrasting logits with and without the prompt) helps in the first round and fades.
- Split-softmax, the authors' method: rescale attention so the system-prompt share becomes π(t)^k, 0 ≤ k ≤ 1, keeping ratios within the prompt and within the history. It gives equal or better stability for a given MMLU loss and works best early. Every method trades some general performance for stability.
Limits
- Instructions, not people. The personas are simple stipulations (language, format, a character trait, a fact), scored by exact functions. Nothing tests a rich persona, memories or values.
- Self-chat between two copies of one model, not humans; eight rounds; 2023–24 models (LLaMA2, GPT-3.5).
- The attention-decay link is shown as co-occurrence plus an idealised theory; the causal claim rests on split-softmax helping.
What it means for Kurisutina
- A slot held in the prompt fades turn by turn. If the person is carried as a system prompt at the start of a long exchange, attention to it drops at every turn the other party speaks. The replica slides toward the base model within a handful of rounds, which is the user's concern made measurable at the instruction level.
- The replica converges on the person it is talking to. The agent took on its interlocutor's instruction over time. People also accommodate to others, but by their own amount; a replica that absorbs its conversation partner at the model's rate is drifting, not carrying on. This is a concrete item for the drift test: run the replica against interlocutors with different styles and check that it converges only as much as the person would.
- The cheap fix is re-injection. Re-inserting the slot before each turn was the most stable option at long horizons. The current harness already rebuilds the full prompt per call (snapshot, reflections, prompt), so it re-injects by construction. A carrying-on replica in live conversation would need the same discipline, at a cost in context. Attention-level fixes need open weights.
- Stability is measurable with deterministic probes. The protocol (probe at every round, score with a function, not a judge) transfers directly: probe a running replica with fixed items whose person-specific answers are known, and track agreement over turns and sessions.
Cross-references
- Lu et al. 2026 (
summaries/carry_on/lu2026_assistant_axis.md) cite this paper as the prompt-level account of drift and extend it to the model's default persona. - Peng et al. 2026 (
summaries/carry_on/peng2025_funhouse.md): the static counterpart, persona data in context that the base model still outweighs.