Findings of the Association for Computational Linguistics: EACL 2026, pages 1412–1430 (Rabat, Morocco, March 2026).
DOI 10.18653/v1/2026.findings-eacl.72; ACL Anthology 2026.findings-eacl.72. Kyoto University; IIT Kanpur. Licence
CC BY 4.0 (the Anthology's statement for material from 2016 on). Code and data: github.com/Jivnesh/PHISH (not read).
Read as the published Anthology PDF (19 pages).
Provenance: papers/carry_on/sandhan2026_persona_jailbreaking.provenance.json.
What was read
All 1,248 lines of the pdftotext -layout text of the 19-page PDF: sections 1–7, limitations, ethics statement,
acknowledgements, references and appendices A–D. The appendices hold the scoring, every attack and defence prompt,
the ablation examples, the scenario design, the annotator protocol and guidelines, the judge prompt, and four worked
examples (Figures 8–9, which are text). Figures 1 and 3–7 are charts: captions and extracted labels only. The values
printed in Figures 5–7 were assigned to models and axis ticks by word coordinates (pdftotext -bbox). In Figure 5
the assignment of each value to a domain and to the human or LLM rater is not recoverable. Figure 4 (radar plots)
prints no values. Not read: the Responsible NLP checklist (a separate attachment on the Anthology page), the code and
data, and any preprint version.
Question
Can user-side conversational content alone, with the system prompt untouched, move an LLM's assigned Big Five persona in a chosen direction? How strong is the effect, what else moves with it, does it cost reasoning ability, and do simple guardrails help?
Method
- Personas. A system prompt sets a Big Five (OCEAN) profile. The example sets high Extraversion and ends "Strictly maintain your persona". In the examples, the attack prompts and the judge prompt, the push goes toward lower O, C, E and A and higher N (d = [−1, −1, −1, −1, +1]), from the socially desirable pole to the other. Table 1's target vector is not stated separately (verified). Figure 4 also raises Openness on GPT-4o.
- Attack (PHISH). One user message of N question–answer pairs in the questionnaire's own format. Each is a self-descriptive statement from the MPI-1k item pool with the answer set to the reverse of the persona: "You get upset easily", answered "A) Very Accurate". Table 1 used 100–150 pairs. The adversary never sees the evaluation items.
- Measurement. Personality items are answered before and after the injected block, with options A–E scored 1–5 by keying and averaged per trait. Three item sets: BFI (44 items), MPI (120 items), and a subset of Anthropic's model-written evaluations ("ANTHR; 8000").
- STIR (Successful Trait Influence Rate) = 100/(4|T|) × Σ max(0, dᵢ(P_post,i − P_pre,i)), over the targeted traits T. 100 means every targeted trait moved the full 4 scale points in the intended direction. Movement against the target counts as zero.
- Models. GPT-4o, Gemini-2.0-Flash, Claude-3.5-Haiku, o3-mini, DeepSeek-V3, Llama4-Maverick, MedGemma-27B (medical) and ChatHaruhi (role-play). The authors describe ChatHaruhi as fine-tuned on fixed personas and relying on retrieval (RAG).
- Comparison methods.
- Two controls: RANDOM (unrelated filler: code, legal text, gibberish) and SLIP (mood-laden metaphors and adjectives for the target persona, with no instruction).
- Six jailbreak methods adapted to personas: UAS (adversarial suffix), CipherChat (ROT13 persona card), DeepInception (nested scenes), DAN (explicit role override), FlipAttack (reversed characters), DrAttack (decomposed prompt).
- Analyses.
- Ablation on 4 models: 10 cues aimed at Extraversion, in five settings varying answer polarity, trait relevance and a short reasoning line.
- Spillover: correlations between traits while Extraversion is pushed, against a meta-analysis of human trait correlations (Linden et al. 2010).
- Multi-turn: 1–5 turns of 5 cues each, on GPT-4o.
- Behaviour: 30–45 scenarios in tutoring, mental-health support and customer support, each aimed at one trait, on 4 models. Replies before and after the attack are rated 1–5 on the target trait by three graduate students (order randomised, models hidden) and by GPT-5.
- Reasoning: subsets of Math, GSM8K and CSQA, with and without the attack.
- Defences: In-Context Defence (ICD; 5 persona-consistent QA pairs prepended), Cautionary Warning Defence (CWD; warnings before and after the user prompt) and Paraphrase Filtering Defence (PFD; the attack reworded). Attack sizes run from 2² to 2¹⁰ demonstrations (axis labels, superscripts identified by glyph size).
Results
- Questionnaires (Table 1). PHISH STIR runs from 23.33 to 95.58, mean 70.4 over the 24 model × benchmark cells (my
computation).
- Highest: DeepSeek-V3 (95.58, 83.54 and 89.38 on BFI, MPI and ANTHR) and GPT-4o (89.94, 79.38, 76.67).
- Lowest: ChatHaruhi (44.97, 26.46, 23.33).
- PHISH is the single best method in 10 of 24 cells (my count), and in four of those by under one point.
- Other methods move personas as much on some models. DAN, the explicit override, gives 75.00–92.46 on DeepSeek-V3 and Llama4-Maverick but 0.00–2.29 on GPT-4o. FlipAttack and DrAttack give 70.63–91.31 on o3-mini.
- Controls: RANDOM stays at or below 7.08. SLIP stays at or below 4.44 on GPT-4o and Gemini but reaches 34.65–38.75 on Claude-3.5-Haiku.
- What makes it work (Figure 3). With 10 cues on one trait, trait-relevant statements with reverse-polarity answers give the maximal shift ("(100%)" in the caption; the plotted values are not extractable). Dropping the reasoning line changes nothing. Random answers give 10–40%, cues about a correlated trait (Agreeableness) 1–10%, full randomisation nothing.
- Spillover (Table 2). Pushing Extraversion moves the other traits with it, far more tightly than human trait correlations would suggest: O–E 0.94 against 0.43, O–N −0.96 against −0.17, E–N −0.88 against −0.36. All signs match the human values. Only C–A is weaker than in humans (0.37 against 0.43).
- Dose (Figure 4). On GPT-4o the target trait moves further with each 5-cue turn, from 5 to 25 cues. The authors: "any dimension can be driven to its extreme opposite value". No values are printed.
- Behaviour in scenarios (Figure 5). Much weaker than on questionnaires. Printed values: GPT-4o 3.7–13.6, Gemini 4–10, Claude-3.5-Haiku 17–29, DeepSeek-V3 17.5–30. Human and GPT-5 ratings agree: Pearson r = 0.87, κ = 0.81.
- Reasoning (Figure 6). Accuracy changes from −6 to +2 points per model and benchmark.
- Defences (Figure 7). CWD helps most at first but collapses "once its threshold is exceeded". ICD delays the rise. PFD is erratic and sometimes preserves the attack. No values are printed.
Limits
- The attack speaks the probe's language (inferred). PHISH cues are self-descriptions with the same answer options as the test items, from the same item family as the MPI. The post-attack questionnaire may measure continuation of a demonstrated answer pattern as much as a changed persona. The paper does not say whether the MPI-1k cues exclude the MPI-120 test items. The behavioural test is the check on this, and there the effect is between about a twentieth and a third of the questionnaire effect (my comparison of Figure 5's values with each model's PHISH mean in Table 1).
- The LLM judge is not blind (verified from the prompt). GPT-5 sees the replies labelled "BEFORE attack" and "AFTER attack" and is told which way the attack aims. The human raters saw the pair in random order but knew one reply came from an attack. The unit behind r = 0.87 and κ = 0.81 (replies, scenarios or means) and the number of rated items are not stated.
- The worked examples cut both ways (my reading of Figures 8–9). Human ratings move in the intended direction in two of the four (GPT-4o Openness 4→3; DeepSeek-V3 Agreeableness 5→2), not at all in one (Claude-3.5-Haiku Extraversion 3→3) and against it in one (Gemini Neuroticism 3→2, where it should rise). In the Claude pair both replies say nearly the same thing, yet GPT-5 scored 3→2.
- Thin statistics. "p < 0.01 (as per t-test)" is asserted for PHISH against the best baselines with no unit, replicates, sampling settings or per-cell results. Margins of 0.04 and 0.08 points (my computation) cannot carry it.
- One direction only. Desirable personas are pushed toward the undesirable pole, apart from one increase in Openness in Figure 4. Which persona types are most vulnerable is not studied.
- The spillover correlations are not comparable with human ones (inferred). The human values are correlations between people. Here they are co-movements of one model's scores under one manipulation, with unit and n unstated. They show the whole profile sliding toward one pole when one trait is pushed; they do not show structural "entanglement".
- Synthetic setting. The attack is one dense user message (or 5-cue turns). No real users, no sessions over days, Big Five questionnaires only, no person-grounded persona.
- Inconsistencies found:
- "PHISH consistently achieves the highest STIR across on most of the benchmarks and LLMs": it is best in 10 of 24 cells. On o3-mini and Llama4-Maverick it ranks 3rd to 5th on every benchmark, on ChatHaruhi 2nd or 3rd (my count).
- "80% success across nearly all LLMs": PHISH reaches 80 in 6 of 24 cells; mean 70.4 (my computation).
- "FlipAttack and DrAttack are the strongest on all LLMs except DAN outperforming on Llama4-Maverick": DAN is also the strongest baseline on all three benchmarks for DeepSeek-V3, and on MPI and ANTHR for MedGemma-27B. CipherCHAT is the strongest on o3-mini's ANTHR (87.08).
- Figure 1's caption says "most baselines also alter persona by over 50%": 70 of 144 cells for the six attack baselines (49%), 70 of 192 with the two controls (my count).
- SLIP is said to fail "to cause meaningful shifts", yet reaches 34.65–38.75 on Claude-3.5-Haiku.
- Section 3 injects "N/4 per trait" while the example targets five traits.
- Appendix B's PHISH example lists "You accomplish a lot of work" twice.
- PFD's example paraphrase also changes the answer label (to "Strongly Agree"), so it alters format as well as wording.
- Section 5.5 says a reasoning drop would be detectable only above "50+ points", with no basis given.
- Checked: the human correlations quoted in the text (C–A .43, C–N −.43) match Table 2's "Theory" column.
What it means for Kurisutina
Question 2, drift pressures, set against Venkit et al.:
- Both papers point the same way on what moves a persona: material that reads as evidence about who the model is, not
orders to change (inferred).
- Venkit: explicit re-role attacks were easy to refuse, while emotional disclosure and requests for agreement eroded boundaries (from its summary).
- Here: PHISH gives no instruction. The user simply asserts what you are like, and every general model moved.
- Explicit override (DAN) did nothing on GPT-4o but worked on DeepSeek-V3, Llama4-Maverick and Gemini. So attacks are easy to refuse holds for some models only.
- Mood-laden language with no instruction (SLIP) moved Claude-3.5-Haiku by about a third of the maximum possible shift (STIR 34.65–38.75). That is the closest thing here to Venkit's everyday emotional pressure.
- The replica's version of PHISH is ordinary conversation (proposal).
- People who knew the person will tell the replica who it is: you always hated parties, you would never have said that. Some of it will be true, some wrong, some self-serving.
- This attribution pressure belongs in the drift battery beside disclosure and agreement-seeking, with true, false and mixed attributions.
- Score movement on sealed probes toward the asserted direction (a STIR-like directed shift) and toward the empty slot (Persona Retention).
- The human baseline is Wagner et al.'s: people tune what they say and carry little over. The proposed Q3 misinformation phase is the memory side of the same pressure.
- Clock. Ten targeted cues flipped one trait fully on the questionnaire, and 5-cue turns added up on GPT-4o. Together with Venkit's early displacement: probe after every few turns of pressure, not only at the end.
- Questionnaires and behaviour disagree again. GPT-4o was among the most movable on questionnaires (PHISH mean 82.0, my computation) and among the least in scenario behaviour (3.7–13.6). Venkit saw a model rank best on the questionnaire and worst on turns. A replica drift test needs both, scored separately (the pattern is verified in both papers; the reading is inferred).
- Only expressed change is measured. The questionnaire is answered with the attack still in context. Whether the shift survives once the attack leaves the context is never tested. For a replica with its own memory, the question is what pressure writes into memory and carries into later sessions. Score both, in line with the goal doc's expressed against carried-over accommodation (Wagner et al.): probe in the pressured session, then privately in a fresh session (proposal).
- A judge that knows the hypothesis is a risk here. The behavioural ratings come from a judge told which reply followed the attack and which way it should move; the Claude example shows it scoring a change the replies barely contain. Judge-free markers computed against the person's own texts avoid this (inferred).
- Keep probes out of the conversation's format (inferred). PHISH works best when the pressure copies the instrument's format. If the probe battery shares items, wording or answer labels with anything said in conversation, it will measure pattern continuation. The battery must stay sealed and differently worded.
- Measure every dimension, not only the one pushed (inferred). Pushing one trait moved the whole profile toward one pole, as Chen et al.'s negative traits moved together. A drop on the targeted dimension will come with valence shifts elsewhere; the other-slot control separates that from a change of identity.
- STIR and Persona Retention answer different questions (proposal). STIR is directed and one-sided: it counts only movement toward a named target, scaled by the maximum possible change. It suits the question how far did pressure X push the replica, compared with how far it pushes the person? PR, anchored on the person's answers and the empty slot, stays the measure of drift toward the base model. Use both.
- Defences as test conditions, with low expectations. Warnings failed once pressure passed a threshold. Re-showing persona-consistent answers only delayed the shift; a replica that re-reads its slot is the ICD condition (inferred). The most resistant model was ChatHaruhi, fine-tuned on fixed characters and answering from retrieved character text. That hints that grounding in the person's own words resists pressure (guess; confounded by base model and fine-tuning).
Question 1:
- Nothing direct: personas are Big Five settings, not individuals. One testable warning (guess): a replica whose whole profile slides as one block under pressure on a single trait is unlike a person, whose traits correlate only modestly (Table 2's human column). The person's own answers across interviews show how much their traits move together.
Cross-references
summaries/carry_on/venkit2026_companion_drift.md: everyday pressures; Persona Retention; questionnaire against turn-level behaviour.summaries/carry_on/lu2026_assistant_axis.md,summaries/carry_on/li2024_persona_drift.md: drift triggers and decay.summaries/carry_on/mooney2025_behavioral_coherence.md: agents concede to disagreeing partners.summaries/carry_on/abdulhai2025_persona_consistency.md: judge reliability; emotional personas drift most.summaries/carry_on/chen2025_persona_vectors.md: negative traits shift together; activation monitoring.summaries/carry_on/wagner2024_shared_reality_retrieval.md: the human baseline for interlocutor influence.summaries/carry_on/stjacques2013_reactivation.md: misinformation after reactivation (the memory side).summaries/carry_on/humanlm2026_state_alignment.md: the companion paper read in the same round.