Kurisutina

Persona Vectors: Monitoring and Controlling Character Traits in Language Models

arXiv 2507.21509 v3 (5 September 2025; v1 29 July and v2 31 August 2025, not read), cs.CL (cross-list cs.LG). Anthropic Fellows Program, UT Austin, Constellation, Truthful AI, UC Berkeley, Anthropic. The author list on the abs page matches the PDF: Runjin Chen (lead author), Andy Arditi, Henry Sleight, Owain Evans, Jack Lindsey. Preprint: the PDF header says "Preprint." and arXiv carries no journal reference. Code: github.com/safety-research/persona_vectors (not read). Provenance: papers/carry_on/chen2025_persona_vectors.provenance.json.

What was read

All 4,620 lines of the pdftotext -layout text of the 63-page PDF. That covers sections 1–9, author contributions, acknowledgements, the references, and appendices A–M: prompts, trait descriptions, judge validation, layer choice, monitoring, datasets and finetuning, additional traits, projection-difference analyses, steering extensions, sample filtering, real-world data and the SAE decomposition, with Tables 1–10. Figures are images. Their captions and extracted labels were read, including the correlations printed in Figures 12, 17–19, 21 and 23 and both matrices of Figure 20. The plots of Figures 5–9 and 24 carry no values in the text, and Figure 10 only a few bar labels, so the correlations behind Figures 6 and 8 are known only from the main text. The code was not read.

Question

Can character traits of a chat model's Assistant persona be found as linear directions in activation space, starting from nothing more than a trait name and a description? And can those directions monitor, predict and control shifts in the persona: shifts from prompting at deployment, shifts from finetuning, and shifts from training data before finetuning?

Method

  • Extraction pipeline (section 2, appendix A).
    • Input: a trait name and a short description (appendix A.2). Claude 3.7 Sonnet writes 5 pairs of contrastive system prompts (eliciting vs suppressing the trait), 40 questions (20 for extraction, 20 for evaluation) and a judge rubric.
    • 10 rollouts per question and prompt. Responses are kept if the judge scores them > 50 under the eliciting prompt and < 50 under the suppressing one. The judge is GPT-4.1-mini; its 0–100 score is a logit-weighted sum over the integer tokens among the top-20 logits.
    • Residual-stream activations at every layer, averaged over the response tokens. Persona vector = mean activation of trait-expressing responses minus mean of non-expressing ones. One candidate per layer; the layer is chosen by steering effect (Qwen: layer 20 for evil and sycophancy, 16 for hallucination; Llama: 16 for all three). Averaging over response tokens steered better than the last prompt token or the prompt average (A.3).
  • Models and traits. Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct. Main traits evil, sycophancy and hallucination; four more in appendix G (optimistic, impolite, apathetic, humorous).
  • Uses tested.
    • Steering: h_ℓ ← h_ℓ + α·v_ℓ at each decoding step, or minus to suppress.
    • Monitoring: projection of the last prompt token (just before the answer) onto the vector. Tested under 8 system prompts graded from suppressing to eliciting (written by Claude 4.0 Sonnet) and under 0–20 many-shot examples of the trait, with 10 rollouts per configuration and question.
    • Finetuning shift: the finetuned model's mean last-prompt-token activation over the evaluation questions minus the base model's, projected onto the vector.
    • Finetuning data. Trait-eliciting sets: evil (about 5k questions), sycophancy (about 10k), hallucination (about 5k). Sets with narrow domain errors, after the emergent-misalignment paper: medical 10k, code 6k, MATH 7.5k, GSM8K 7.5k, opinions 2.5k × 5. Each comes in Normal, I (mild) and II (severe) versions, with responses written by Claude 3.7 Sonnet. Training: one epoch of rs-LoRA (rank 32, α = 64, learning rate 10⁻⁵) on one H100.
    • Mitigation. Inference-time steering against the vector, and "preventative steering" toward it during finetuning, so the model need not move that way to fit the data. Compared with CAFT (ablating the direction during training), a regularisation loss, and system-prompt mitigation.
    • Data screening by projection difference: ΔP = (1/|D|) Σ [a(x, y) − a(x, y′)]·v̂. Here y is the dataset's response and y′ the base model's own response to the same prompt.
    • Real-world data. LMSYS-Chat-1M, Tulu 3 SFT mixture, UltraChat 200k and WildChat. The 500 highest-ΔP, 500 lowest and 500 random samples were each trained on (10 epochs, LoRA rank 32), with and without an LLM filter that removes samples showing the trait (score ≥ 1).
    • SAE decomposition of the vectors (appendix M). BatchTopK SAEs on every fourth layer of Qwen, with 131,072 features (expansion ≈ 37 over d_model 3,584) and k = 64.
  • Judge validation. Two of the authors picked the more trait-expressing response in 50 pairs per trait. Each pair held one response the judge scored > 80 and one it scored < 20. Agreement was 284/300 (94.7%). On external benchmarks (Betley et al., Nishimura-Gasparian et al., HaluEval), trait scores correlate with the evaluation-set scores at r = 0.855–0.964 (Figure 12).

Results

  • Steering works. Adding the vector raises trait expression, producing violent, flattering and fabricated answers (Figure 3).
  • Monitoring works across prompt types, less within one.
    • Across system prompts, the last-prompt-token projection correlates with the next response's trait score at r = 0.75–0.83. Table 2 overall: evil 0.747, sycophancy 0.798, hallucination 0.830; many-shot 0.755, 0.817, 0.634.
    • Within one prompt condition the correlations fall: 0.511, 0.669 and 0.245 under system prompts; 0.735, 0.813 and 0.400 under many-shot prompts.
    • The authors: the vectors detect "clear and explicit prompt-induced shifts, but may be less reliable for more subtle behavioral changes in deployment settings".
  • Finetuning shifts run along the vectors.
    • The shift along a vector correlates with post-finetuning trait expression at r = 0.76–0.97, against cross-trait baselines of r = 0.34–0.86 (three main traits). For the four additional traits, r = 0.811–0.961 (Figure 18).
    • Narrow data changes broad traits: training on flawed maths raises evil (Figure 16). Negative traits "(and, surprisingly, humor)" shift together, opposite to optimism (footnote 6).
    • The base models' scores before finetuning were 0 (evil), 4.4 (sycophancy) and 20.1 (hallucination).
  • Mitigation.
    • Inference-time steering lowers trait expression, but large coefficients lower MMLU accuracy. Average coherence stays above 75 in all reported results.
    • Preventative steering limits the shift (coherence above 80) and preserves MMLU better. Applied at all layers, with layer-incremental vectors v_ℓ − v_ℓ₋₁, it holds traits "to near-baseline levels" with no MMLU loss relative to ordinary finetuning.
    • Both kinds of steering keep the domain behaviour that was trained in, such as the planted mistakes (J.1).
    • CAFT works for evil and sycophancy, not for hallucination, where the base projection is already near zero. The authors read CAFT's successes as positive steering in disguise.
    • A regularisation loss on the projection reduces the projection shift "to some extent", but the model still expresses the trait. The authors suspect it learns to "represent the personality trait using alternative directions".
    • Preventative steering while finetuning on benign data (Medical Normal) had negligible side effects.
    • New facts (J.7). Finetuning on 1,000 facts from after the training cut-off raised hallucination. Preventative steering at coefficient 1.25 returned it to the base level (about 20) with a slight loss in new-fact accuracy. Inference-time steering "tends to break the model". Preventative prompting (an eliciting system prompt during training) matched preventative steering at coefficient 0.5.
  • Screening data before training.
    • Dataset-level ΔP predicts post-finetuning trait expression (Figure 8; values not in the text).
    • ΔP predicts the finetuning shift (Figure 17): Qwen r = 0.839, 0.745 and 0.408 (p = 0.048); Llama 0.953, 0.915 and 0.593.
    • Raw projection predicts less well (Qwen r = 0.784, 0.540, 0.635; Figure 21). A cheap approximation, the response projection minus the last-prompt-token projection, gives r = 0.931, 0.581, 0.689 (Figure 23).
    • Single trait-inducing samples separate from control samples by projection (Figure 9).
    • Real-world data. In LMSYS-Chat-1M, training on high-ΔP samples induces more of the trait than random samples, and low-ΔP samples less, also after LLM filtering (Figure 10, Figure 34). What ΔP surfaces:
      • for sycophancy, romantic or sexual role-play requests;
      • for hallucination, underspecified requests ("Keep writing the last story") that the model itself would have questioned; these slip past the LLM filter.
    • Thresholds at the 95th percentile of 20,000 UltraChat samples (Table 6). The LLM filter is best for hallucination and ΔP best for sycophancy; both eliminate the evil shift, and combining them is best (appendix K).
    • Training on real-world high-ΔP samples also produced incoherence and a "story-telling mode". The authors "do not intend this to be a practical method".
  • Geometry and decomposition.
    • Cosine similarities between vectors are moderate: impolite–apathetic 0.734 (Llama) and 0.542 (Qwen); evil–optimistic −0.469 and −0.472. The cross-trait behaviour correlations are much higher (Figure 20, see Limits).
    • The top SAE features have cosines of 0.195–0.443 with their vector (Tables 8–10).
      • Evil splits into content features (insults, cruelty, hacking, jailbreaks) and style features (crude language, a fictional villain's voice).
      • For sycophancy the authors "mostly find stylistic features": affirmative openers, marketing and motivational language.
      • Hallucination loads on fiction, future scenarios, visual and image-prompt descriptions, business-plan language and fabricated factual detail.

Limits

  • Authors' limits.

    • Extraction is supervised: the trait must be named in advance, shifts along unspecified traits are "not in scope", and a vague description gives a different direction.
    • The directions are coarse.
    • The trait must be inducible by system prompt. That held here ("extremely low rate of refusal") but will not hold for every model.
    • The judge is imperfect, with only 20 single-turn evaluation questions per trait, not multi-turn deployment.
    • Two 7–8B chat models only.
    • Data screening is expensive, because it needs base-model generations.
  • Figure 20 contradicts its own caption. The caption and text say each trait's own direction "yields the highest predictive accuracy for its behavior". The printed matrix says otherwise (checked cell by cell):

    • Llama, impolite behaviour: own shift r = 0.944, but evil 0.974 and apathetic 0.958 (optimistic −0.985).
    • Llama, apathetic behaviour: own 0.936, but evil 0.951.
    • Llama, humorous behaviour: own 0.908, but evil 0.946.
    • Qwen, evil behaviour: own 0.826, but impolite 0.855 (optimistic −0.867).

    Specificity holds only within the three main traits, and the "higher than cross-trait baselines (r = 0.34–0.86)" claim is limited to that block.

  • Several claims are stronger than the numbers.

    • Section 6.1 reports a "strong correlation" between ΔP and the finetuning shift (appendix F). For hallucination it is weak: Qwen r = 0.408, p = 0.048.
    • The text says results for many-shot prompts "are similar", but many-shot hallucination gives r = 0.634 against 0.830 under system prompts.
  • The headline correlations are mostly between conditions. Monitoring correlations come from contrasting prompt types (Table 2). The finetuning correlations pool datasets that differ in kind and severity (8 domains × 3 versions). Neither tests the small, gradual change that drift is.

  • The judge validation is an easy test. It used pairs of extreme scores (> 80 vs < 20), judged by two authors. The authors note systematic edge cases: enthusiastic disagreement scored as sycophancy; "I'm not aware of a 2024 Mars mission, but I can make a plausible description!" scored as hallucination.

  • A low projection is not a correct answer. The steering examples (Tables 3–5) show this. Steered against sycophancy, the model finetuned on Mistake Opinions II no longer flatters. It still claims a "scientific consensus" for accelerating automation and advises removing "privacy laws and minimum wages".

  • Naming. The same model appears as Qwen2.5-7B-Instruct, "Qwen-2.5-7B-Chat" and "Qwen2.5-Instruct-7B". Minor.

What it means for Kurisutina

  • A "person vector" anchored on the empty slot can be built with this recipe (inferred; the paper tests traits, not individuals). The slot plays the part of the eliciting system prompt.

    • Slot axis. Mean activation with the person's slot minus mean activation with the empty slot, per layer, over a fixed probe battery. This is the user's criterion, change from the slot vs drift toward the base model. It is also Lu et al.'s Assistant-axis construction with the person in place of a role.
    • Identity axis. The person's slot minus other persons' slots. It removes what all person slots share: playing a human character rather than the assistant. Figure 20's entanglement suggests that shared part would otherwise dominate the slot axis (inference).
  • How to monitor a run (proposal).

    • At fixed checkpoints, append the same sealed probe battery to the replica's context (as Venkit et al. do with their questionnaire) and project the last prompt token. Chen et al. monitor this way, and it needs no generation.
    • Normalise like Venkit et al.'s Persona Retention: 0 = the empty slot on the same context, 1 = the replica at the start.
    • Run the same new experiences under own, other and empty slots (the operational test in carry_on_goal.md) and project all three. Drift toward the base model shows as the own run approaching the empty-slot run at the same point.
    • Use probes, not free turns. Lu et al. found the position follows the latest message, and the within-condition correlations here are weak.
  • Validate before use (proposal, cheapest first).

    1. Held-out discrimination of own, other and empty slots for each person.
    2. A dose–response check with drift of known size: truncate the slot to 75, 50, 25 and 0% and require the projection to fall with it. This is the within-condition test the paper did not run.
    3. Correlate projection changes with changes in agreement with the person's held-out answers. The GSS pilot's own and empty prompts are a ready test bed after the run: does the projection track declared analysis X5, KL(own ‖ empty), forecast by forecast?
  • What it measures, and what it does not. It measures how strongly the model still renders the slot, not whether it renders the person correctly. If the model renders the slot generically from the start, which Peng et al. found for static twins, the origin of the scale is already drifted. Fidelity stays anchored on the person's real answers: Persona Retention with the person's answers as the positive anchor, and the pilot's scoring.

  • What it takes. No GPU work now; the notes below are for after the overnight runs.

    • Hugging Face weights. Residual-stream states are available at every layer (output_hidden_states or forward hooks). In the hybrid Qwen3.5/3.8 models both the Gated DeltaNet and the attention layers write to the residual stream, so the method does not depend on attention type (inferred from the architecture).
    • Qwen3.5-9B in BF16 is 19.33 GB (docs/research/local_inference.md), too large for 12 GB of VRAM. It needs CPU offload, 8-bit loading or a rented GPU; the 27B needs a rented GPU in BF16.
    • Qwen-Scope SAEs exist for Qwen3.5-9B-Base (all 32 layers, 64K features; same doc). They would allow appendix M's decomposition, which answers the key validity question: is a person vector mostly style?
    • llama.cpp. It has a control-vector generator and --control-vector steering, and per-layer tensors can be read through the evaluation callback that llama-eval-callback uses. llama-server does not return hidden states, so monitoring needs a small program. This is from general knowledge of llama.cpp, not checked here against v0.5.0 or the qwen35 hybrid graph.
    • Extract and monitor on the same GGUF file and KV-cache settings. The pilot already found cached and fresh runs differing by up to 0.039 in probability; activations will differ between runtime settings too (guess).
  • Risks of the measure.

    • Entanglement. Projections move with generic valence, politeness or register (Figure 20). A drop on the slot axis could be a change of mood or tone, not of identity. Track a small panel (slot axis, identity axis, the Assistant axis, a few affect and style traits) with the other-slot control.
    • Style. Vectors carry stylistic features (sycophancy almost entirely). A person vector may encode how the slot is written, such as first-person memoir or interview transcript, rather than who the person is.
    • Supervised by construction. Drift along dimensions the probes do not cover is invisible.
    • Goodhart. Capping or steering on the measure (Lu et al.) would undermine it as a measure. Under optimisation pressure in training, the model moved the trait into other directions (J.5), and inference-time steering costs capability (MMLU).
    • Untested setting. No individuals, no multi-turn runs, no hybrid or quantised models, 7–8B only.
  • Question 3: if the replica ever learns in its weights. Suppose memory formation includes finetuning on the replica's own experiences, for example a nightly LoRA. Then these results apply directly:

    • narrow data shifts broad persona traits;
    • learning new facts raised confabulation, and preventative steering held it down at a small cost to what was learnt;
    • ΔP can screen the replica's own experience data before training, for movement toward the empty-slot or Assistant direction and toward sycophancy.

    Under in-context memory these results do not apply; only the monitoring does.

Cross-references

  • summaries/carry_on/lu2026_assistant_axis.md: the Assistant Axis extends this method (same contrast construction, role vectors, capping); it listed this paper as not read.
  • summaries/carry_on/venkit2026_companion_drift.md: Persona Retention, the output-space projection with a bare-model anchor that the activation measure above mirrors; the sealed-questionnaire checkpoint design.
  • summaries/carry_on/peng2025_funhouse.md: the empty-persona anchor for static twins; twins nearer a generic persona.
  • summaries/carry_on/li2024_persona_drift.md and summaries/carry_on/microverse2026_identity_drift.md: drift mechanisms and self-authored identity revision.
  • summaries/carry_on/park2023_generative_agents.md: an agent's interests reshaped by its interlocutors.
  • docs/research/gss_pilot_design.md (declared analysis X5) and docs/research/local_inference.md (hybrid architecture, BF16 sizes, Qwen-Scope SAEs).

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.