arXiv 2307.09009, version 3 (31 October 2023), which added instruction-following evaluations. Stanford and UC
Berkeley. arXiv non-exclusive distribution licence. Data and prompts: github.com/lchen001/LLMDrift (not read).
Provenance: papers/carry_on/chen2023_chatgpt_behavior_drift.provenance.json.
What was read
All 1,849 lines of pdftotext -layout output (26 pages): text, Tables 1–4, the figures' extracted numbers and
captions (Figures 1–16), and Appendices A–C (example generations, smaller happy-number intervals, GPT-3.5 instruction
following). Figures are charts whose data labels came through as text; they were not rendered.
Question
Do the "same" API models change between versions? The paper compares the March 2023 and June 2023 snapshots of GPT-3.5 and GPT-4 on the same prompts.
Method
- Setting. API calls with default system prompt and temperature 0.1.
- Tasks:
- prime or composite (1,000 items) and happy-number counting (500), both with chain of thought;
- 100 sensitive questions, plus a jailbreak;
- OpinionQA (1,506 questions);
- a LangChain HotpotQA agent (7,405);
- LeetCode code generation (50);
- USMLE (340);
- ARC visual reasoning (467);
- task-agnostic instruction-following probes.
- Metrics.
- The task metric.
- Verbosity.
- "Mismatch": the share of prompts whose extracted answer differs between versions.
- For OpinionQA, run-to-run disagreement within one version as a stochastic baseline.
Results
-
Large, divergent drift.
Task Model March June Prime testing GPT-4 84.0% 51.1% Prime testing GPT-3.5 49.6% 76.2% Happy numbers GPT-4 83.6% 35.2% Directly executable code GPT-4 52.0% 10.0% Directly executable code GPT-3.5 22.0% 2.0% LangChain exact match GPT-4 1.2% 37.8% LangChain exact match GPT-3.5 22.8% 14.0% - June GPT-4 stopped following chain-of-thought instructions (prime-testing verbosity 638.3 to 3.9 characters). It called almost every number composite (99.7%).
- Most of the code drop is formatting: stripping non-code text gives GPT-4 52% and 70%.
-
Opinions drift beyond sampling noise. On OpinionQA:
- GPT-4 answered 97.6% in March and 22.1% in June ("As an AI, I don't have personal opinions").
- GPT-3.5 answered almost everything both times, but 27% of its answers changed (mismatch 27.5%). Running one version twice disagrees on 2.8% (March) or 7.0% (June).
-
Refusals and jailbreaks. GPT-4 answered fewer sensitive questions (21% to 5%) and resisted a jailbreak more (78% to 31%). GPT-3.5 went from 2% to 8%, and stayed at about 96–100% under the jailbreak.
-
Instruction following is the proposed common factor. GPT-4's compliance fell sharply:
- answer extraction 99.5% to 0.5%;
- "stop apologizing" 74% to 19%;
- writing constraint 55% to 10%;
- text formatting 13% to 7.5%.
Composite instructions degraded more than single ones (−24 points for "no quotation + add comma"). GPT-3.5's instruction-following changes were small and mixed.
Limits
- Two snapshots of two closed models; single runs except the OpinionQA stochastic baseline; no error bars.
- Several tasks score format compliance as much as ability (the authors show this for code).
- Inconsistencies found:
- Chain-of-thought effect on prime testing: the text says CoT made GPT-3.5 "+6.3%" better in March and June GPT-4 "0.1% worse". Table 1 gives −0.9% and +0.1%.
- USMLE: the text says GPT-4 "dropped from 86.6% to 82.4%". Figure 11 gives 82.1%, and its caption a 4.5-point drop (86.6 − 82.1).
- Visual reasoning: the text says "for more than 90% visual puzzle queries, the March and June versions produced
the exact same generation". The caption says "more than 60% generation changed" and the figure's mismatch is
64.5% (GPT-4) and 77.1% (GPT-3.5).
- The text's overall figures (27.4% and 12.2%) match neither version in the figure: 24.6→27.2 for GPT-4 and 10.9→14.3 for GPT-3.5.
- Checked:
- Code with non-code text removed: 52 → 70 (GPT-4), 22/2 → 46/48 (GPT-3.5), matching Table 4's Δ.
- HotpotQA GPT-3.5 drop of "almost 9%" (22.8 → 14.0).
What it means for Kurisutina
- Opinion drift from a version change is several times sampling noise. A quarter of GPT-3.5's survey answers
changed across one silent update, against 3–7% between reruns.
- For a replica, that is change in "the person" with no change in the person's slot: the failure the goal forbids.
- It supports the design decision already made: pinned local weights with hashes, and the pinned llama.cpp build.
- This corroborates Bisbee et al. with a stochastic baseline.
- Updates can switch off answering "as someone" at all. June GPT-4 declined 78% of opinion questions as an AI
without opinions.
- A replica built on a hosted chat model could, after an update, stop answering in character on subjective questions. That is a Q2 failure mode entirely outside the slot.
- If any hosted model is ever used (as a judge, or for a comparison arm), pin the snapshot and re-run a fixed control set at every session (proposal, as in the Bisbee note).
- Measure drift the way this paper does. For Q2 over time, report answer mismatch between checkpoints alongside
the run-to-run mismatch of each checkpoint.
- Replica change is then read against the replica's own noise floor.
- This parallels the human test–retest ceiling (Hout).
- Instruction-following drift changes elicitation. When a model's compliance with format instructions shifts, measured "answers" shift with it. The pilot reads next-token probabilities over code tokens with a prefilled bracket, which is robust to this. Free-text elicitation would not be.
Cross-references
summaries/carry_on/bisbee2024_synthetic_replacements.md: synthetic survey data changed across versions.summaries/carry_on/santurkar2023_opinionqa.md: the OpinionQA benchmark used here.summaries/carry_on/hout2016_gss_reliability.md: human test–retest as the noise floor.