Kurisutina

How is ChatGPT's behavior changing over time?

arXiv 2307.09009, version 3 (31 October 2023), which added instruction-following evaluations. Stanford and UC Berkeley. arXiv non-exclusive distribution licence. Data and prompts: github.com/lchen001/LLMDrift (not read). Provenance: papers/carry_on/chen2023_chatgpt_behavior_drift.provenance.json.

What was read

All 1,849 lines of pdftotext -layout output (26 pages): text, Tables 1–4, the figures' extracted numbers and captions (Figures 1–16), and Appendices A–C (example generations, smaller happy-number intervals, GPT-3.5 instruction following). Figures are charts whose data labels came through as text; they were not rendered.

Question

Do the "same" API models change between versions? The paper compares the March 2023 and June 2023 snapshots of GPT-3.5 and GPT-4 on the same prompts.

Method

  • Setting. API calls with default system prompt and temperature 0.1.
  • Tasks:
    • prime or composite (1,000 items) and happy-number counting (500), both with chain of thought;
    • 100 sensitive questions, plus a jailbreak;
    • OpinionQA (1,506 questions);
    • a LangChain HotpotQA agent (7,405);
    • LeetCode code generation (50);
    • USMLE (340);
    • ARC visual reasoning (467);
    • task-agnostic instruction-following probes.
  • Metrics.
    • The task metric.
    • Verbosity.
    • "Mismatch": the share of prompts whose extracted answer differs between versions.
    • For OpinionQA, run-to-run disagreement within one version as a stochastic baseline.

Results

  • Large, divergent drift.

    Task Model March June
    Prime testing GPT-4 84.0% 51.1%
    Prime testing GPT-3.5 49.6% 76.2%
    Happy numbers GPT-4 83.6% 35.2%
    Directly executable code GPT-4 52.0% 10.0%
    Directly executable code GPT-3.5 22.0% 2.0%
    LangChain exact match GPT-4 1.2% 37.8%
    LangChain exact match GPT-3.5 22.8% 14.0%
    • June GPT-4 stopped following chain-of-thought instructions (prime-testing verbosity 638.3 to 3.9 characters). It called almost every number composite (99.7%).
    • Most of the code drop is formatting: stripping non-code text gives GPT-4 52% and 70%.
  • Opinions drift beyond sampling noise. On OpinionQA:

    • GPT-4 answered 97.6% in March and 22.1% in June ("As an AI, I don't have personal opinions").
    • GPT-3.5 answered almost everything both times, but 27% of its answers changed (mismatch 27.5%). Running one version twice disagrees on 2.8% (March) or 7.0% (June).
  • Refusals and jailbreaks. GPT-4 answered fewer sensitive questions (21% to 5%) and resisted a jailbreak more (78% to 31%). GPT-3.5 went from 2% to 8%, and stayed at about 96–100% under the jailbreak.

  • Instruction following is the proposed common factor. GPT-4's compliance fell sharply:

    • answer extraction 99.5% to 0.5%;
    • "stop apologizing" 74% to 19%;
    • writing constraint 55% to 10%;
    • text formatting 13% to 7.5%.

    Composite instructions degraded more than single ones (−24 points for "no quotation + add comma"). GPT-3.5's instruction-following changes were small and mixed.

Limits

  • Two snapshots of two closed models; single runs except the OpinionQA stochastic baseline; no error bars.
  • Several tasks score format compliance as much as ability (the authors show this for code).
  • Inconsistencies found:
    • Chain-of-thought effect on prime testing: the text says CoT made GPT-3.5 "+6.3%" better in March and June GPT-4 "0.1% worse". Table 1 gives −0.9% and +0.1%.
    • USMLE: the text says GPT-4 "dropped from 86.6% to 82.4%". Figure 11 gives 82.1%, and its caption a 4.5-point drop (86.6 − 82.1).
    • Visual reasoning: the text says "for more than 90% visual puzzle queries, the March and June versions produced the exact same generation". The caption says "more than 60% generation changed" and the figure's mismatch is 64.5% (GPT-4) and 77.1% (GPT-3.5).
      • The text's overall figures (27.4% and 12.2%) match neither version in the figure: 24.6→27.2 for GPT-4 and 10.9→14.3 for GPT-3.5.
  • Checked:
    • Code with non-code text removed: 52 → 70 (GPT-4), 22/2 → 46/48 (GPT-3.5), matching Table 4's Δ.
    • HotpotQA GPT-3.5 drop of "almost 9%" (22.8 → 14.0).

What it means for Kurisutina

  • Opinion drift from a version change is several times sampling noise. A quarter of GPT-3.5's survey answers changed across one silent update, against 3–7% between reruns.
    • For a replica, that is change in "the person" with no change in the person's slot: the failure the goal forbids.
    • It supports the design decision already made: pinned local weights with hashes, and the pinned llama.cpp build.
    • This corroborates Bisbee et al. with a stochastic baseline.
  • Updates can switch off answering "as someone" at all. June GPT-4 declined 78% of opinion questions as an AI without opinions.
    • A replica built on a hosted chat model could, after an update, stop answering in character on subjective questions. That is a Q2 failure mode entirely outside the slot.
    • If any hosted model is ever used (as a judge, or for a comparison arm), pin the snapshot and re-run a fixed control set at every session (proposal, as in the Bisbee note).
  • Measure drift the way this paper does. For Q2 over time, report answer mismatch between checkpoints alongside the run-to-run mismatch of each checkpoint.
    • Replica change is then read against the replica's own noise floor.
    • This parallels the human test–retest ceiling (Hout).
  • Instruction-following drift changes elicitation. When a model's compliance with format instructions shifts, measured "answers" shift with it. The pilot reads next-token probabilities over code tokens with a prefilled bracket, which is robust to this. Free-text elicitation would not be.

Cross-references

  • summaries/carry_on/bisbee2024_synthetic_replacements.md: synthetic survey data changed across versions.
  • summaries/carry_on/santurkar2023_opinionqa.md: the OpinionQA benchmark used here.
  • summaries/carry_on/hout2016_gss_reliability.md: human test–retest as the noise floor.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.