Robert Geirhos and Kristof Meding (joint first authors), Felix A. Wichmann; University of Tübingen. NeurIPS 2020; arXiv:2006.16736v3 (18 December 2020). Code: github.com/wichmann-lab/error-consistency (not inspected). Read by researcher R6 (capability), 6 October 2026, as the method source for comparing a replica's errors with a person's. Provenance: papers/capability/geirhos2020_error_consistency.provenance.json.
What was read
- Read in full: the text of the pdftotext conversion (10,843 lines, 26 pages, most of them plot labels): abstract, sections 1–4, broader impact, references, and supplement S.1–S.6 (terminology, derivation of bounds, the simulation of confidence intervals, accuracy and κ, model details, shape bias).
- Not read: plot-only lines (bare numbers, scatter markers, confusion-matrix cells in Figure SF.8) and Table 1's per-observer accuracies, which extracted as unlabelled numbers. Every caption and all prose were read.
What they did
- Question: two decision makers (people, networks, or a mix) can have the same accuracy with different strategies. How can one tell whether they fail on the same items?
- Method ("error consistency"). For observers i and j with accuracies pᵢ and pⱼ on the same n trials:
- observed overlap c_obs = share of trials where both are right or both are wrong;
- chance overlap for independent observers, c_exp = pᵢpⱼ + (1 − pᵢ)(1 − pⱼ);
- κ = (c_obs − c_exp) / (1 − c_exp) (Cohen's kappa repurposed). κ = 0 is chance-level, κ > 0 shared strategy, κ < 0 opposite.
- Analytical bounds on κ given c_exp, and confidence intervals from simulating 88.2M independent-observer experiments.
- The textbook confidence interval for kappa is shown to be wrong and still widely used.
- Data:
- 10 human observers (from Geirhos et al. 2019) classifying 224-px images shown for 200 ms into 16 categories. Stimuli: cue-conflict (1,280 trials per observer), edge drawings (160) and silhouettes (160). Plus 2 observers on plain ImageNet images.
- Compared with 16 ImageNet-trained PyTorch CNNs and the recurrent "brain-like" CORnet-S.
- Procedure note: shuffle presentation order per participant, since fatigue and serial dependence can create "trivial" consistency.
Main results (verified)
- Humans agree with each other well beyond chance on which items are hard: human–human κ = 0.33 ± 0.02 (cue conflict), 0.32 ± 0.04 (edges), 0.48 ± 0.03 (silhouettes), about 0.37 (ImageNet).
- CNNs agree with each other even more, across architectures: CNN–CNN κ about 0.45–0.67 by experiment. The highest pair is κ = 0.793 (DenseNet-121 against ResNet-18, different families and depths). This holds although humans were more accurate: silhouettes, human 0.75 against CNN 0.54.
- CNN–human consistency is near chance for cue-conflict and edge stimuli.
- It does not improve with ImageNet accuracy (R² = 0.001 and 0.003). For silhouettes it rises with accuracy (R² = 0.253); for ImageNet images it falls (R² = 0.214).
- "AlexNet from 2012 is just as error-consistent as recent models."
- The recurrent "current best model of the primate ventral stream" behaves like a feedforward ResNet-50.
- Cue conflict: CORnet-S–human κ = .066, ResNet-50–human .068, AlexNet .080, human–human .331.
- CORnet-S against ResNet-50 κ = .711.
- Recurrence alone did not make behaviour human-like.
- Models with more shape bias (trained on stylised images) were more error-consistent with humans on cue-conflict stimuli.
- The authors' principles:
- "Making consistent errors on a trial-by-trial basis is a necessary condition" for similar strategies, though not a sufficient one.
- "Human-level accuracy does not imply human-like decision making."
- Errors are the informative trials when both systems are near ceiling, hence the name.
- A metric a model was optimised for (here Brain-Score) can look good while an unoptimised behavioural metric exposes the gap.
Limits
- The chance model treats each observer's correct and wrong trials as independent of the item (a binomial with fixed accuracy). It does not account for item difficulty. κ between two humans therefore mixes shared item difficulty with shared strategy. Both are "beyond chance" here.
- Binary right/wrong only: which wrong answer is ignored (confusion matrices are shown separately). Small numbers of trials per condition (160) make κ noisy near ceiling, where its denominator shrinks.
- Vision with 10 observers, rapid presentation; no repeat sessions of the same observer, so no within-person test–retest κ.
What it means for Kurisutina (inference)
- The core fidelity metric for a capability profile is accuracy-corrected item agreement. Compare replica and person on the same items, by κ on right/wrong, and preferably on the specific wrong answer. Matching the person's score is necessary and nowhere near sufficient.
- Calibrate against human references:
- person against own retest: the ceiling. Not in this paper; it needs a repeat session.
- person against other people of similar level: the population floor, about 0.3–0.5 here.
- replica against the person must beat replica against other people (individuation) and approach the retest ceiling.
- Because human–human κ already contains shared item difficulty, person-specific fidelity is the excess over what an item-difficulty model predicts. That is where item response theory at person level comes in (my extension, not in the paper).
- A warning for a trained replica: networks converge on a shared machine strategy regardless of architecture (κ up to 0.79 between families), distinct from humans. A replica fine-tuned on general data will tend to share the machine pattern. Only training on the person's own item-level behaviour moves it (compare Maia,
mcilroyyoung2020_maia.md). - Recurrence and capability settings do not by themselves produce human-like errors (CORnet-S). Error patterns must be tested directly, not assumed from architecture.
- Use the corrected confidence intervals (simulation or the corrected formula), and shuffle item order across sessions to avoid fatigue-driven artefacts.