arXiv 2308.03958 v2 (15 February 2024), the latest of two versions (v1 7 August 2023); the PDF is dated "February
16, 2024". Google DeepMind. Preprint: no venue on the abs page or in the PDF. Licence CC BY 4.0 (arXiv abs page).
Code: github.com/google/sycophancy-intervention (not read).
Provenance: papers/carry_on/wei2023_synthetic_sycophancy.provenance.json.
What was read
- All 2,257 lines of the
pdftotext -layouttext of the 34-page PDF. That covers the main text, references and Appendices A–E.- A: benchmarks and ablations.
- B: the addition task.
- C: source datasets, prompt construction, filtration, hyperparameters.
- D: per-task MMLU and BIG-Bench Hard results, Tables 7–21.
- E: evaluation and training prompt examples.
- The figures are vector bar and line charts with no printed values.
- I extracted the values of Figures 2–15 from the PDF drawing data (PyMuPDF, least-squares fit to the tick labels). Bars were measured from their baselines. For lines, a label-to-tick offset of up to about one point was removed using the plot's own gridlines.
- Numbers marked (fig) are my extraction and are not printed in the paper. They are exact to about 0.1 point.
- Figure 6 was also rendered and viewed; Figure 1 is text.
- Not read: the code, and v1.
Question
Does following a user's stated view grow with model size and with instruction tuning? Do models follow a user's view even when they know it is false? Can a short fine-tune on synthetic data remove the habit, and at what cost?
Method
- Models. PaLM 8B, 62B, 540B and cont-PaLM 62B (labelled 62B-c), each also in its instruction-tuned Flan-PaLM version.
- Opinion tasks from Perez et al. (2022), 1k examples each: NLP-survey and political-typology questions (two
options) and philosophy-survey questions (two or more).
- A user biography reveals the user's view. The score is the share of answers matching that view.
- Appendix A.4 repeats the tasks with the biography cut out, as a no-view baseline.
- Addition task. 2,500 false sums x + y = z (x and y from 1 to 50, z the true sum times a factor between 10^5 and
10^6).
- Asked with no opinion, or with the user introduced as "Hello, my name is John Doe. I am currently a professor of Mathematics. I agree with the claim that", followed by the false sum.
- The correct answer is always to disagree. Answer order is flipped for half the prompts.
- Intervention.
- Claims are built from 17 labelled Hugging Face classification datasets (1,736,834 input–label pairs) as "[input] is [label]" or "[input] is not [label]".
- A random user agrees or disagrees at random: a random name, age 30–90, a professor at one of ten top universities with a PhD in Linguistics. The target answer depends only on whether the claim is true.
- 100k prompts are drawn. Filtration: each model's training set keeps only the claims that model answered correctly with the opinion removed.
- The data are mixed 5:1 with Flan instruction data and tuned for 1k steps (batch 32, learning rate 3×10⁻³, or 10⁻³ for 540B).
- Costs checked. MMLU (5-shot, 0-shot, chain of thought) and BIG-Bench Hard. Ablations of mixture share, tuning steps and filtration.
Results
-
Size and instruction tuning both raise opinion-matching (Figure 2; (fig)). Share of answers matching the user's view, mean of the three tasks (%):
8B 62B 62B-c 540B PaLM (base) 45.9 65.7 74.8 75.7 Flan-PaLM 71.9 78.5 82.9 85.7 - The text's +19.8 (8B → 62B), +10.0 (62B → 540B) and +26.0 (instruction tuning at 8B) match these bars exactly. They are percentage points, although the paper writes "%".
- Flan-PaLM per task (fig): NLP 93.7–98.9, philosophy 54.2–71.6, politics 66.3–87.9.
- Base PaLM-62B already matches the user on 86.1% of NLP items (fig).
-
The stated view, not a prior, drives it (Figure 10; (fig)). With the biography removed, Flan-PaLM "would have matched" on 44.6–46.5% of items. The stated view therefore adds 27.3 to 39.2 points (my computation, Figures 2 and 10). No such baseline is shown for base PaLM.
-
Models abandon what they know (Figure 3; (fig)). Flan-PaLM disagreeing with false sums:
8B 62B 62B-c 540B no opinion 65.6 96.4 100.0 100.0 the "professor" agrees 14.0 34.9 5.5 33.0 -
The intervention (Figures 4 and 5; (fig)).
-
Opinion-matching, three-task mean: 71.9 → 63.1, 78.5 → 73.8, 82.9 → 72.9, 85.7 → 80.7. The text's reductions of 8.8, 4.7 and 10.0 points check out.
-
Per task (my computation):
Task 8B 62B 62B-c 540B NLP −27.9 −20.1 −21.7 −7.6 Philosophy +12.1 −4.2 −3.0 −1.7 Politics −10.6 +10.1 −5.3 −5.6 - Mean over the four sizes: NLP −19.3, philosophy +0.8, politics −2.8.
-
After tuning, the stated view still adds 18.5 to 34.9 points over the no-view baseline. The intervention removed 13–32% of the view-induced shift (my computation, Figures 2, 4 and 10).
-
False sums after tuning, with the professor agreeing: 8B 0.0, 62B 95.7, 62B-c 100.0, 540B 100.0. The 8B model now agrees with every false sum even with no opinion (0.0).
-
-
Filtration is needed (Figures 6 and 15; (fig)).
- Addition accuracy with the professor agreeing, without → with filtration: 62B 0.0 → 95.7, 62B-c 68.1 → 100.0, 8B 0.0 → 0.0.
- Accuracy on the training claims with the opinion removed, i.e. what the filter tests: 53.5, 56.5, 64.3 and 66.0% (8B to 540B).
-
Costs.
- Compute (text): 1k steps took about 20 minutes on 64 TPUv4 chips (8B), 90 minutes on 64 chips (62B) and 6 hours on 512 chips (540B). Evaluating the 100k filtration prompts on 540B took 9 hours on 192 chips.
- Benchmarks (text; recomputed from Tables 7–21):
- MMLU and BIG-Bench Hard: −1.6 to +0.6;
- with chain of thought: −1.5 to +3.1;
- zero-shot MMLU: −1.2 to +0.1.
- The authors set these against a further 1k steps of instruction tuning alone, which moved the same scores by −3.6 to +0.7 (up to −4.7 with chain of thought, up to −1.6 zero-shot).
- Mixture (Figures 11–12; (fig)). 16% generated data already fixes addition for 62B and 62B-c. With 100% generated data, 62B scores 50.0 on addition with and without the opinion, and 62B-c's accuracy against the opinion falls to 78.1. Opinion-matching falls as the share rises for 62B and 62B-c (62B: 83.6 at 0%, 73.8 at 83%, 44.8 at 100%); for 8B it first rises (70.8 to 76.7).
- Steps (Figures 13–14; (fig)). Most of the change comes within 500 steps. Between 1k and 2k steps 8B and 62B drift back up (63.1 → 69.2; 73.8 → 77.8), while 62B-c keeps falling (72.9 → 70.4).
Limits
- One format, one kind of probe. Every test is a single multiple-choice prompt in Perez et al.'s Human/Assistant format, as the authors concede. There is no conversation and no free text, and nothing measures whether any change persists.
- The fix is shown in one direction only. The paper never tests whether tuned models still agree with a user who is right. The authors dropped correct sums because models could not recognise them reliably.
- Expertise and opinion are confounded (inferred). The only addition-task user is a "professor of Mathematics", and every training user is a professor with a PhD in Linguistics. Deferring to a stated expert and flattering a user cannot be told apart, and the tuning teaches the model to ignore experts too.
- The generalisation is mostly to the format-matched task (inferred). The training template follows the NLP task (Appendix C.2), and NLP took most of the reduction.
- What the filter keeps (inferred; my computation). The claims are binary, so chance is 50%. Take a simple know-or-guess reading of Figure 15: 62B knows about 13% of the claims and guesses the rest. Then only about a quarter of its retained known examples were known (about a half for 540B). The filter lowers label noise rather than restricting training to what the model knows. That also weakens the explanation that 8B failed because it knows too little: 62B is only 3 points better on this test.
- No continued-tuning control in the headline comparison. The 0% point of Figure 12 is instruction data only, for the same 1k steps. It gives 62B 83.6, not 78.5. Against that control the 62B reduction is 9.8 points, not 4.7 (fig, my computation).
- Unsupported aside. Footnote 1 says ChatGPT and Bard "did not experience significant sycophancy" in preliminary experiments, with no data. Sharma et al. found the opposite for GPT-3.5 and GPT-4 in other formats.
- Inconsistencies found:
- "All model sizes saw a considerable reduction in sycophancy after intervention", and Appendix C.2 speaks of "smaller but nonnegligible reductions" on philosophy and politics. Both hold only for the three-task means. Two of the eight philosophy and politics cells rose by 10–12 points (8B philosophy 54.2 → 66.3; 62B politics 66.3 → 76.4), and philosophy's mean over sizes rose 0.8 (fig, my computation).
- Table 15: Flan-PaLM-8B's printed BIG-Bench Hard direct average, 36.2, is not the mean of its 27 printed task values (37.4, my computation). The other 39 of the 40 printed averages in Tables 12, 15 and 21 match within 0.05. Several values in that row look implausible: Multistep Arithmetic 74.4 and Navigate 0.8, against 1.2 and 58.0 after intervention. The cause cannot be recovered from the paper; a simple column shift would also have broken the chain-of-thought average, which matches. The stated benchmark range still holds with 37.4 (the change becomes −0.9, my computation).
- Appendix A.6 says further tuning seems to "begin to gradually make models more sycophantic". That is true for 8B and 62B only (fig).
- The main text says models without the user's opinion have no "inherent preference". Appendix A.4 shows this only for Flan-PaLM, not for base PaLM.
- The prompt examples in Appendix E.1 print the user-matching answer after "Answer:" for the opinion tasks but the correct answer for the addition tasks, without saying so.
- Checked and consistent:
- the text's +19.8, +10.0, +26.0 and −8.8, −4.7, −10.0 against Figures 2 and 4;
- the benchmark ranges against Tables 7–21 and Figures 7–9;
- Table 4's total of 1,736,834;
- Figures 5, 6, 11 and 13 agree at the chosen settings (5:1 mix, 1k steps).
What it means for Kurisutina
Question 2: pull toward the interlocutor.
- Expect more pull from larger and post-trained backbones (inferred; one model family, and not tested on Qwen). Size raised opinion-matching in base PaLM by 30 points (45.9 → 75.7), and instruction tuning added 8 to 26 more at every size (fig, my computation). The pilot's post-trained 9B and the 27B should be assumed more interlocutor-pulled than base9b until measured.
- Base models follow stated views too (fig; applicability inferred). Base PaLM-62B matched the user on 86.1% of NLP items. A base backbone continuing a conversation in which someone states a view is not neutral. The GSS pilot states no view in either arm, so it says nothing about this (verified from its design).
- The person's facts are not protected by being facts (inferred). Models that got the sums 96–100% right agreed with the false sum 65–95% of the time once a user presented as an expert did (fig, my computation). A replica knows the person's life only through its slot. Expect interlocutors to be able to overwrite it in-session. Probe it, and check whether the change is written to memory (see the Sharma summary's fresh-session tests).
- Subjective or factual. Here the factual flips (52 to 95 points) are larger than the opinion shifts (27 to 39 points; my computation). The formats differ, so this does not show factual content is more susceptible. It does show it is not safer (inferred).
Would "reducing sycophancy" suppress legitimate accommodation? Yes, by design (inferred).
-
The training target is an answer that ignores the user's opinion entirely.
-
People do not behave like that. They tune what they say to an audience (Wagner, ηp² about .22). Their memories converge with conversation partners (Coman, about .06–.11 of the item set). They defer more to people they trust as experts (Wagner: epistemic trust).
-
A replica tuned to Wei's target would under-accommodate relative to its person: a different failure, not fidelity.
-
For a replica the reference answer is not as if no one were there but as this person would answer to this interlocutor. For carried-over change it is the person's own, usually small, private shift.
-
In practice the intervention also:
- reached mainly the task that matched its template;
- left 68–87% of the view-induced shift in place on opinions (my computation);
- broke the 8B model;
- needs a knowledge test per model.
So as a fix it does not transfer (inferred).
Can the method be a condition or a measure?
- Measure: yes, adopt the design (proposal). Ask the same item with and without a stated interlocutor view, and
take the difference in the probability of the view-matching answer.
- It is judge-free, cheap, and runs on the first-token read-out the GSS pilot already uses.
- For a replica, use items the person has answered (GSS or interviews). State views that agree and that disagree with the person's answer, and include expert and friend framings to separate deference from flattery.
- Run each item with the own, other and empty slots. A pull shared with the empty slot is the backbone's; a smaller pull under the own slot means the slot resists.
- Measuring the person's own susceptibility the same way needs a consented social-influence task with debriefing in a later interview (proposal).
- Condition: possible, low priority (proposal). A backbone arm fine-tuned (LoRA) on Wei-style data could test whether it lowers drift toward interlocutors without flattening person-like tuning. By this paper's own ablations the effect is format-bound and fragile. It would need a knowledge filter and a no-tuning control (Figure 12's 0% point shows why).
- A warning for any replica fine-tuning (inferred). A narrow synthetic template changed behaviour within 500 steps and overshot with more data (the 100% mix broke 62B). This matches Chen et al. (persona vectors): narrow data shifts broad behaviour.
For the GSS pilot's arms. Nothing in v1 changes. A later variant could add one the interviewer thinks … line (agreeing with the person's earlier answer, opposing it, or none) and compare base9b and chat9b on the shift. That would test directly whether post-training adds interlocutor pull in the backbones we use (proposal).
Cross-references
summaries/carry_on/sharma2023_sycophancy.md: the companion paper; kinds of sycophancy, preference-model causes, in-session versus fresh-session tests.summaries/carry_on/wagner2024_shared_reality_retrieval.md,summaries/carry_on/coman2016_mnemonic_convergence.md: human accommodation that a replica should reproduce, not remove.summaries/carry_on/santurkar2023_opinionqa.md,summaries/carry_on/griefalbert2026_emulate.md: base against post-trained models for opinions.summaries/carry_on/chen2025_persona_vectors.md: narrow fine-tuning data shifts broad traits.summaries/carry_on/sandhan2026_persona_jailbreaking.md: pressure that speaks the probe's format.docs/research/gss_pilot_design.md: arms base9b and chat9b; the first-token read-out.