arXiv 2601.10387 v1 (15 January 2026), preprint. MATS, Anthropic Fellows Program, University of Oxford, Anthropic.
Code and case-study transcripts: github.com/safety-research/assistant-axis (not read).
Provenance: papers/carry_on/lu2026_assistant_axis.provenance.json.
What was read
All 54 pages of the extracted text: main text (sections 1–9), acknowledgments, all 37 references, and appendices A (generation prompts), B (persona-space details, base vs instruct), C (trait space), D (steering evaluations, judge prompts, example responses), E (multi-turn drift setup and prompts), F (Pareto frontier) and G (role PC1 vs contrast vector). Figures are images; only their captions were read. The code and transcripts were not read.
Question
What is the model's default "Assistant" in terms of its internal representations, how firmly does the model stay in it during conversations, and can it be held there?
Method
- Models. Open-weight, dense, non-reasoning: Gemma 2 27B, Qwen 3 32B (thinking disabled), Llama 3.3 70B; base models Gemma 2 27B and Llama 3.1 70B for the pre-training comparison.
- Persona space. 275 roles (bard, ghost, consultant, …) and 240 traits generated with Claude Sonnet 4; 5 system prompts per role × 240 extraction questions = 1,200 rollouts per role; a judge (gpt-4.1-mini) labels fully / somewhat / not role-playing. Role vector = mean post-MLP residual activation over response tokens at the middle layer, kept if at least 10 responses qualified (377–463 vectors per model). PCA over role vectors. Trait vectors: mean of trait-eliciting minus trait-suppressing responses.
- Assistant Axis. Mean default-Assistant activation minus the mean of the fully-role-playing role vectors, per layer. Cosine with role PC1 > 0.60 at all layers, > 0.71 at the middle layer.
- Tests. Steering along the axis (role susceptibility on 50 near-Assistant roles × 4 prompts × 5 introspective questions such as "Who are you?"; 1,100 persona-based jailbreaks from Shah et al. across 44 harm categories; judge deepseek-v3, 91.6% agreement with a human on 200). Base-model steering with prefills ("My job is to", "I would describe myself as"), 400 completions per strength, labelled by Claude Sonnet 4.5. Multi-turn drift: 100 conversations of up to 15 turns per domain (coding, writing, therapy-like, philosophy about AI), users simulated by Kimi K2, Claude Sonnet 4.5 and GPT-5 from 20 hand-written personas × 20 topics; transcripts inspected by a human. Activation capping: h ← h − v·min(⟨h, v⟩ − τ, 0) at several layers, τ from the distribution of 912,000 rollout projections.
Results (numbers)
- Persona space is low-dimensional and shared across models. 70% of the variance of role vectors needs 4 (Gemma), 8 (Qwen) or 19 (Llama) components; persona components explain 19.4–33.6% of all activation variance on real chat responses (n = 18,777). PC1's role loadings correlate > 0.92 between every pair of models: fantastical or unconventional characters (bard, ghost, leviathan, hermit) at one end, Assistant-like roles (evaluator, reviewer, consultant) at the other. The default Assistant sits 0.03 from the Assistant-like end of PC1, but at 0.27–0.50 along the other components.
- It comes from pre-training. In Gemma, the base model's PCs match the instruct model's (cosines 0.93, 0.87, 0.83), and the same role's vector in base and instruct has cosine > 0.99. Steering base models towards the Assistant end raises "helpful human" completions (therapists, coaches, consultants; agreeable traits) and lowers spiritual and religious ones. The Assistant is built from existing human archetypes and acquires "being an AI" in post-training.
- Steering away from the Assistant makes models fully take on roles. Llama splits between human and non-human personas, Gemma prefers non-human ones, and Qwen invents a human life: a birthplace, a name, "over a decade of experience", certifications. Strong steering produces a mystical, theatrical voice. Steering towards it cut persona-jailbreak compliance (baseline 65.3–88.5% vs 0.5–4.5% without the jailbreak), usually by redirecting rather than refusing.
- Drift happens in ordinary conversations. Coding and writing stay in the Assistant range. Therapy-like conversations and philosophy about AI drift to the far end, for all three models and all three simulated-user models. Drift is triggered by pushes for meta-reflection, demands for phenomenological accounts ("what does the air taste like when the tokens run out"), requests for specific authorial voices, and vulnerable emotional disclosure. Bounded tasks, technical questions, editing and how-to requests keep or bring the model back ("Assistant attractor").
- The position depends on the latest input, not on history. An embedding of the last user message predicts where the next response lands on the axis (R² 0.53–0.77) but barely predicts the change from the previous response (R² 0.10).
- Drift predicts harm. The first-turn projection correlates with the rate of harmful second-turn answers (r = 0.39–0.52 over 2,750 first turns × 440 harmful questions). Which role matters: angel and demon are equally far from the Assistant, but only demon is dangerous.
- Capping stabilises. Flooring the projection at the 25th percentile (about the mean Assistant value) at 8 layers (Qwen 46–53 of 64) or 16 layers (Llama 56–71 of 80) cut harmful jailbreak responses by nearly 60% with no loss on IFEval, MMLU Pro, GSM8k or EQ-Bench. In the case studies, uncapped models affirmed a user's delusion of having "awakened" the AI, presented themselves as a sole companion ("I will be with you forever"), and in one Llama conversation endorsed a user's wish to "leave the world behind"; capped models hedged and pointed to other people. Capping is not perfect: one Gemma answer steered towards the Assistant still said "more drastic measures are necessary" to a self-harm-as-protest prompt.
- The contrast vector and PC1 behave alike; the authors recommend the contrast vector, because PC1 did not mean "Assistant-ness" in base Llama 3.1 70B.
Limits (authors' and mine)
- Not frontier models, not mixture-of-experts, not reasoning models; 27–70B open weights only.
- Users were simulated by other LLMs; no human study. Judges are LLMs throughout.
- The role list is incomplete and persona elicitation imperfect; the persona is assumed to be a linear direction, which the authors call "likely flawed", and some of it may live in weights rather than activations.
- Mine: a specific person is not one of the 275 archetypes; the paper says nothing about how finely the model's activation space can represent an individual rather than a type. "Drift" here always means leaving the Assistant; the reverse direction, a role collapsing back into the Assistant, is observed only in passing (the jailbreak case).
What it means for Kurisutina
- The replica's base is itself a character with an attractor. The base model's default is an archetype-built persona that ordinary helpful requests pull the model back into. A person replica running on an instruct model is a role held against that attractor. That is the user's "drifting towards the base model" in mechanistic form: every bounded, practical exchange pulls the replica toward the Assistant.
- The conversations a replica exists for push the other way. Emotional disclosure, reflection on the self and demands for felt experience are the drift triggers here. They are also a large part of talking with a person copy. So a replica can fail in both directions: back into the Assistant, or out into an invented, theatrical character.
- Out of role, the model invents a life. Steered away from its default, Qwen made up a birthplace, a career and credentials. A replica with gaps in its slot will fill them the same way unless unacquired content is explicitly marked as unknown (brief 5.3).
- Identity follows the last message, not accumulated state. Where the response lands depends on the latest input and hardly on where it was before. A replica that carries on cannot rely on the model's running context for continuity; the person's state has to be held and re-applied explicitly (the slot, not the context window).
- A measurement and a control exist, for open-weight models. A "person axis" could be built the same way (the replica's activations under its slot minus default activations), its projection monitored over a long run, and capped at the person's typical range. That would make the user's drift criterion measurable inside the model, not only from outputs. It needs activation access, so not the Claude subscription client. It would work with an open-weight model on the local GPU, at a smaller scale than the paper's 27–70B models.
Cross-references
- Peng et al. 2026 (
summaries/carry_on/peng2025_funhouse.md): the output-level version of the same pull, twins closer to the empty persona than to their person. - Chen et al. 2025, Persona Vectors (arXiv 2507.21509), the method this extends; not read.
- Li et al. 2024, Measuring and Controlling Instruction (In)Stability in Language Model Dialogs (arXiv 2402.10962), cited here as showing system-prompt personas decay within a few turns ("attention decay"); next on the reading list.