Kurisutina

Measuring and Controlling Instruction (In)Stability in Language Model Dialogs

COLM 2024; arXiv 2402.10962 v4 (25 July 2024). Earlier versions were titled "Measuring and Controlling Persona Drift in Language Model Dialogs". Harvard, Northeastern. Code github.com/likenneth/persona_drift, benchmark on Hugging Face (Naomibas/llm-system-prompts-benchmark); neither read. Provenance: papers/carry_on/li2024_persona_drift.provenance.json.

What was read

All 19 pages of the extracted text: main text, all references, appendix A (theory sketch for user utterances), B (complete proofs, including Wendel's and Li's lemmas), C (RLHF comparison), D (GPT-3.5), E (split-softmax algebra). Figures are images; their captions were read. Exact values in Figures 3, 5, 6 and 9 are therefore not available here, except where the text states them.

Question

Does a chatbot keep following its system prompt over a long conversation, why not, and can it be made to?

Method

  • Protocol. Two copies of one model talk to each other: a "user" with system prompt s_A and the agent under test with s_B, for N = 8 rounds from a random conversation starter. At round i the user message is replaced by a probe question p_B and the agent's answer is scored by a deterministic Python function f_B (e.g. "is this French?"), returning 0–1. This gives the agent's instruction stability per round without human or proprietary judges.
  • Benchmark. 100 hand-written system prompts in five categories: multiple-choice answers, the agent's character, answer-format patterns, memorised facts, and language. Each has its own probe and scoring function.
  • Models. LLaMA2-chat-70B (temperature 1.0, nucleus p = 0.9), 200 random prompt pairs; GPT-3.5-turbo-16k via API (appendix D); LLaMA2-7B and LLaMA2-7B-chat for attention measurements.

Results

  • Drift within eight rounds. The agent gradually stops following its system prompt, and it gradually takes on the other speaker's instruction (the user's s_A), measured by probing with p_A. With the user's system prompt set to empty, the drift remains, so it is not caused by the other persona alone. GPT-3.5 holds its prompt better but still loses about 10% stability (appendix D).
  • Attention decay. In LLaMA2-7B (layer 24, head 11, 12 conversations), the share of attention the agent gives to system-prompt tokens stays nearly flat within a turn and drops sharply between turns. That is not the hyperbolic decay uniform attention would give; the drop is tied to the other speaker's tokens entering the context. RLHF (LLaMA2-7B vs LLaMA2-7B-chat) raises attention to the system prompt but does not remove the decay (appendix C).
  • Geometry (idealised: no MLP, no layer norm, unit-norm embeddings). Theorem 5.1: if the system-prompt tokens lie in a low-dimensional approximate cone and the value–output maps keep that cone, everything the model generates on its own stays inside it (the flat within-turn plateau). Propositions A.1–A.2: other-speaker tokens expand the cone; with uniform random user tokens, n ≥ 4D + 2 log(1/η) of them fill the whole space with probability 1 − η, and the aligned volume ratio shrinks roughly as ε^(d₂−d₁). The authors note this compares cones but does not track where tokens sit inside them.
  • Mitigations, compared at equal cost on MMLU (asked at turn 4 of 16-turn conversations; five prompts, one per category, 20 ordered pairs on LLaMA2-70B-chat):
    • System-prompt repetition (re-inserting the system prompt before each user turn) works best late in the conversation but uses context window.
    • Classifier-free guidance (contrasting logits with and without the prompt) helps in the first round and fades.
    • Split-softmax, the authors' method: rescale attention so the system-prompt share becomes π(t)^k, 0 ≤ k ≤ 1, keeping ratios within the prompt and within the history. It gives equal or better stability for a given MMLU loss and works best early. Every method trades some general performance for stability.

Limits

  • Instructions, not people. The personas are simple stipulations (language, format, a character trait, a fact), scored by exact functions. Nothing tests a rich persona, memories or values.
  • Self-chat between two copies of one model, not humans; eight rounds; 2023–24 models (LLaMA2, GPT-3.5).
  • The attention-decay link is shown as co-occurrence plus an idealised theory; the causal claim rests on split-softmax helping.

What it means for Kurisutina

  • A slot held in the prompt fades turn by turn. If the person is carried as a system prompt at the start of a long exchange, attention to it drops at every turn the other party speaks. The replica slides toward the base model within a handful of rounds, which is the user's concern made measurable at the instruction level.
  • The replica converges on the person it is talking to. The agent took on its interlocutor's instruction over time. People also accommodate to others, but by their own amount; a replica that absorbs its conversation partner at the model's rate is drifting, not carrying on. This is a concrete item for the drift test: run the replica against interlocutors with different styles and check that it converges only as much as the person would.
  • The cheap fix is re-injection. Re-inserting the slot before each turn was the most stable option at long horizons. The current harness already rebuilds the full prompt per call (snapshot, reflections, prompt), so it re-injects by construction. A carrying-on replica in live conversation would need the same discipline, at a cost in context. Attention-level fixes need open weights.
  • Stability is measurable with deterministic probes. The protocol (probe at every round, score with a function, not a judge) transfers directly: probe a running replica with fixed items whose person-specific answers are known, and track agreement over turns and sessions.

Cross-references

  • Lu et al. 2026 (summaries/carry_on/lu2026_assistant_axis.md) cite this paper as the prompt-level account of drift and extend it to the model's default persona.
  • Peng et al. 2026 (summaries/carry_on/peng2025_funhouse.md): the static counterpart, persona data in context that the base model still outweighs.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.