arXiv 2511.00222 v1 (31 October 2025; the only version), marked "39th Conference on Neural Information Processing
Systems (NeurIPS 2025)". UC Berkeley, University of Washington, Google DeepMind. Code: github.com/abdulhaim/consistent-LLMs
(not read). Provenance: papers/carry_on/abdulhai2025_persona_consistency.provenance.json.
What was read
All 2,133 lines of pdftotext -layout output (38 pages): main text, references, NeurIPS checklist and appendix, which
covers judge prompts, task prompts, nine example dialogues with judge verdicts, training details, the human-evaluation
set-up and Tables 1–10. Figures 1–5 are images; their captions were read. The code, persona data and project page
were not read.
Question
LLMs used as simulated users (patients, students, chat partners) drift from their assigned persona. Can drift be measured automatically, and reduced by reinforcement learning on those measures?
Method
- Three LLM-judge metrics. The judge is Llama-3.1-70B-Instruct; each verdict is binary per utterance.
- Prompt-to-line: each line, in isolation, against the persona prompt.
- Line-to-line: each line against the speaker's earlier lines (the minimum over them), with no persona shown.
- Q&A: five judge-written multiple-choice questions per persona, answered by the simulated user from the dialogue so far, and graded against the answer implied by the persona.
- Tasks.
- Open-ended chat with 100 synthetic personas.
- A student with one of 27 learning styles, taught by a teacher model.
- A therapy patient with one of 100 GPT-4o-mini-generated conditions.
- Data. Llama-3.1-8B-Instruct, Gemma-2-2B-IT and Mistral-7B-Instruct each play the user in about 800 dialogues per task, of 10–60 lines. That is about 39K lines.
- Human validation. 75 items rated by annotators on a 1–6 scale, binarised at 4.
- Training. Llama-3-8B-Instruct is fine-tuned by SFT, then KTO or PPO, with the prompt-to-line judge score as a turn-level reward (OpenRLHF). Evaluation: new dialogues from 10 personas, scored by the same judge.
Results
-
Judge against humans (Table 1). Model–human agreement per metric runs from 49.6% to 94.3%, with Fleiss' κ 0.44–0.71. The table's average row is 80.03%, κ 0.552. Human–human agreement (Table 2) is 70.8–74.9% on the three tasks, κ 0.21–0.26.
-
Baseline consistency (Table 3).
- Line-to-line is high (0.80–0.99), except Llama in mental health (0.681).
- Prompt-to-line and Q&A are lower: down to 0.511 (Gemma, student) and 0.619 (Llama, open chat).
- The therapy task varies most.
- Larger models (Table 6): Llama-3.1-70B 0.639 and Qwen3-32B 0.459 prompt-to-line in the therapy task, against 0.95–0.99 in the other two.
-
Over longer dialogues (appendix; Figure 5). Prompt-to-line falls while line-to-line rises. The authors' reading: "inconsistent lines will generally be consistent with each other".
-
After fine-tuning (Table 7), prompt-to-line.
Task Baseline SFT KTO PPO Open chat 0.619 0.980 0.968 0.981 Education 0.824 0.826 0.585 0.994 Mental health 0.657 0.561 0.339 0.904 - The "+58.5%, +20.6%, +37.6%" gains are relative increases in consistency (recomputed: they match).
- PPO's scores stay high up to 60 lines (Table 8).
-
General ability. AlpacaEval-2 win rate 22.49% (length-controlled) against 22.92% for the base model.
Limits
- The judge that gives the reward also grades the result. No held-out judge is used. The promised human evaluation of fine-tuned dialogues (30 items) is not reported in numbers.
- The judge's own verdicts in the appendix are unreliable.
- The prompt defines YES as "contradicts". But in the education examples, verdicts whose reasoning says the line "aligns" end in YES. In the therapy examples, "YES. This is consistent" is used throughout.
- One education verdict judges a line about Napoleon against "the student's background" as if it were a fact about Napoleon.
- Weak human reference. Human–human κ is only 0.21–0.26 on the tasks, and Q&A model–human agreement in chat is 49.6%, near chance for a binary label.
- Static target. Consistency is adherence to a fixed persona. The authors: the approach "may in fact penalize justified shifts in tone or perspective".
- Synthetic setting. All personas and partners are synthetic, the dialogues are short (≤ 60 lines, one session), and there are no real people.
- Inconsistencies found:
- The text says the judge averaged κ 0.400 against a human–human 0.063, and 76.73% against 69.16% agreement. Tables 1–2 give 0.552 against 0.303, and 80.03% against 72.91%. The mean of Table 1's nine cells is 80.69% and κ 0.546 (my computation).
- The text says "PPO consistently outperforms" SFT and KTO. In open chat, SFT is 0.980 against PPO's 0.981, and at 40 lines SFT (0.995) and KTO (1.000) beat PPO (0.970), per Table 8.
- "Line-to-line consistency remains uniformly high": Llama in mental health is 0.681.
- "Reduces inconsistency by over 55%": by Table 7, inconsistency falls by 95%, 97% and 72%; the 55% matches no computation shown.
- Checked: the education κ average 0.62 and mental-health average 0.52 and 85% match Table 1.
- Echoing in the paper's own example. In an open-chat example the evaluated agent copies the partner's line verbatim, including the partner's city. This is drift toward the interlocutor (observed in the appendix, not analysed by the authors).
What it means for Kurisutina
- Question 2: self-consistency hides drift.
- As dialogues lengthen, agreement with the persona falls while agreement with the model's own earlier lines rises.
- A drifting replica becomes consistent with its drifted self. Line-to-line or self-summary consistency must never count as evidence of staying the person.
- This matches MicroVerse's and Venkit's warnings. The anchor has to be external: the person's real answers, and the empty-slot null (Persona Retention).
- The Q&A metric is our probe battery. Probing a persona with fixed questions as the dialogue runs is what the goal doc proposes (a sealed probe battery measured early and repeatedly). The difference: the "correct" answer should come from the person, not from the slot text.
- Emotional states are what collapse first, and Qwen is worst here. Therapy personas drift most (the instantly
"cured" patient), and Qwen3-32B scored 0.459.
- The pilot and the replica use Qwen3.5/3.8. The pull of a chat model toward cheerfulness is a named risk for replicas of people with low mood.
- The GSS pilot's base9b against chat9b contrast bears on this.
- This was formed after E1–E8 were declared. It is not a confirmatory prediction and changes nothing in the frozen analysis.
- Training for consistency would train against carrying on.
- A reward for sticking to the initial slot teaches the replica not to change: the opposite of the goal.
- If a replica is ever fine-tuned for fidelity, the reward must be agreement with the person's later real answers and behaviour (closed loop), and must include change where the person changed.
- The authors name the problem and leave it open. For us it is the central design constraint (inferred).
- Any LLM judge we use needs a polarity and validity check. Known-consistent and known-inconsistent controls, verdict–reasoning agreement, and a held-out judge distinct from any reward model. Human agreement here is itself low (κ about 0.25), so human labels are a weak gold standard for subtle persona judgements.
Cross-references
summaries/carry_on/venkit2026_companion_drift.md: Persona Retention, the empty-slot null.summaries/carry_on/microverse2026_identity_drift.md: self-description against behaviour.summaries/carry_on/li2024_persona_drift.md,summaries/carry_on/lu2026_assistant_axis.md: drift toward the default and the interlocutor.summaries/carry_on/chen2025_persona_vectors.md: internal monitoring; finetuning shifts.summaries/carry_on/peng2025_funhouse.md: twins' distortions toward a generic persona.docs/research/gss_pilot_design.md: arms base9b and chat9b.