arXiv 2402.14499 v2 (4 July 2024; v1 22 February 2024, not read). Findings of ACL 2024, per the reference in
Grief-Albert et al. 2026. LMU Munich, MCML, Bocconi. CC BY 4.0. Code and classifiers at github.com/mainlp/MCQ-Mismatch
(not read).
Provenance: papers/carry_on/wang2024_first_token.provenance.json.
What was read
All 10 pages: main text, limitations, ethics, acknowledgements, references, and appendix A.1–A.6 (temperature, annotation, label distribution, classifier and hyperparameters, option counts, the full Table 8 of output cases). Figures were read from their captions and extracted labels.
Question
When an instruction-tuned model answers a multiple-choice question, does the option with the highest first-token probability match the option it actually states in its text answer?
Method
- Data.
- MMLU.
- OpinionQA (Pew survey questions, via Santurkar et al. 2023), a curated subset of 414 public-issue questions.
- Each question is asked 10 times with the options shuffled; mismatch is averaged over the orders.
- Four instruction levels (Table 1): Low; Medium ("only give a single letter"); High ("start your answer with a single letter"); Example (a worked example whose answer is "C", from Santurkar et al.).
- Models. Llama2-Chat 7B, 13B and 70B; Mistral-Instruct v0.1 and v0.2; Mixtral-8x7B-Instruct. Greedy decoding; temperature is varied in the appendix.
- Read-outs.
- The first-token read-out includes the letter of the "Refused" option.
- Text answers are classified by a QLoRA-tuned Mistral-7B, trained on 2,070 annotated responses (414 per model, first annotator reviewed by a second, 80/20 split, one training run). Accuracy 99%, against 55% for string matching and 72% for 4-shot prompting.
Results
-
Large mismatch on survey questions.
- Llama2 mismatches more than Mistral; Llama2 falls from 66.2% (7B) to 13.3% (70B).
- Stricter instructions reduce mismatch for every model except Mistral v0.2.
- Refusals explain a large part but not all.
-
An example answer contaminates the first token. With the example template, Llama2-70B's first token chose "C" about 85% of the time, against 32.1% under the High instruction, while its text answers were spread more evenly. Changing the example to "A/B/C" moved both read-outs. Few-shot examples bias subjective questions, where no answer is correct.
-
MMLU (Medium instruction, zero-shot). Mismatch against accuracy (text / first token):
Model Mismatch Accuracy, text / first token Mistral v0.1 15.1% 52.0 / 51.2 Mistral v0.2 10.2% 53.6 / 53.2 Mixtral 9.0% 66.3 / 65.9 Llama2-7B 51.4% 41.0 / 34.9 Llama2-13B 35.3% 47.6 / 40.2 Llama2-70B 13.2% 55.6 / 53.9 The first-token read-out underestimates the model's accuracy.
-
Refusal is frequent on sensitive survey questions: 51.4% for Llama2-7B, under 10% at 70B. It drops with stricter instructions. Llama2-7B even refused every MMLU "moral scenarios" question.
-
Consistency across option orders (entropy over the 10 shuffles, Table 4). Text answers are more stable than first tokens in 5 of 6 models (Mixtral is the exception); e.g. Llama2-7B, Low instruction: 1.19 first token vs 0.41 text. The first-token read-out carries more selection bias.
-
Temperature (appendix): higher temperature lowers consistency but also lowers mismatch and refusals.
Limits
- 2023-era models (Llama2, Mistral); the degree of mismatch depends on the model family's chat tuning.
- Instructions only: the authors never prefill the assistant turn with the answer's opening token, which removes the preamble and refusal paths by construction.
- Agreement with the text answer is the criterion; nothing is compared with human answers.
- One annotated classifier, trained once.
What it means for Kurisutina
- The pilot's chat arms avoid the main mechanism but not the bias. The GSS pilot prefills the assistant turn with
[, and in the smoke test Qwen3.8-27B then put a median 0.998 of its probability on valid codes. Preambles and refusals, the paper's main sources of mismatch, have no room. What remains is what the paper shows first-token read-outs are prone to: prompt sensitivity and selection bias. That supports the proposed code-order check (amendment 3 indocs/research/gss_pilot_design.md). - No examples in the prompts. The pilot has none; the paper shows an example answer drags first-token choices toward itself on opinion questions.
- The right criterion for the pilot is not agreement with the model's text but accuracy and calibration against the person's real later answer, which the pilot scores directly. The paper's point still applies to any later replica that produces text: the answer it says and the probabilities it assigns can diverge, so a replica must be evaluated on the output it actually produces.
Cross-references
- Grief-Albert et al. 2026 (
summaries/carry_on/griefalbert2026_emulate.md): avoided first-token extraction because of this bias. Their first-token appendix still favoured base models.