Kurisutina

Beyond accuracy: quantifying trial-by-trial behaviour of CNNs and humans by measuring error consistency

Robert Geirhos and Kristof Meding (joint first authors), Felix A. Wichmann; University of Tübingen. NeurIPS 2020; arXiv:2006.16736v3 (18 December 2020). Code: github.com/wichmann-lab/error-consistency (not inspected). Read by researcher R6 (capability), 6 October 2026, as the method source for comparing a replica's errors with a person's. Provenance: papers/capability/geirhos2020_error_consistency.provenance.json.

What was read

  • Read in full: the text of the pdftotext conversion (10,843 lines, 26 pages, most of them plot labels): abstract, sections 1–4, broader impact, references, and supplement S.1–S.6 (terminology, derivation of bounds, the simulation of confidence intervals, accuracy and κ, model details, shape bias).
  • Not read: plot-only lines (bare numbers, scatter markers, confusion-matrix cells in Figure SF.8) and Table 1's per-observer accuracies, which extracted as unlabelled numbers. Every caption and all prose were read.

What they did

  • Question: two decision makers (people, networks, or a mix) can have the same accuracy with different strategies. How can one tell whether they fail on the same items?
  • Method ("error consistency"). For observers i and j with accuracies pᵢ and pⱼ on the same n trials:
    • observed overlap c_obs = share of trials where both are right or both are wrong;
    • chance overlap for independent observers, c_exp = pᵢpⱼ + (1 − pᵢ)(1 − pⱼ);
    • κ = (c_obs − c_exp) / (1 − c_exp) (Cohen's kappa repurposed). κ = 0 is chance-level, κ > 0 shared strategy, κ < 0 opposite.
    • Analytical bounds on κ given c_exp, and confidence intervals from simulating 88.2M independent-observer experiments.
    • The textbook confidence interval for kappa is shown to be wrong and still widely used.
  • Data:
    • 10 human observers (from Geirhos et al. 2019) classifying 224-px images shown for 200 ms into 16 categories. Stimuli: cue-conflict (1,280 trials per observer), edge drawings (160) and silhouettes (160). Plus 2 observers on plain ImageNet images.
    • Compared with 16 ImageNet-trained PyTorch CNNs and the recurrent "brain-like" CORnet-S.
  • Procedure note: shuffle presentation order per participant, since fatigue and serial dependence can create "trivial" consistency.

Main results (verified)

  • Humans agree with each other well beyond chance on which items are hard: human–human κ = 0.33 ± 0.02 (cue conflict), 0.32 ± 0.04 (edges), 0.48 ± 0.03 (silhouettes), about 0.37 (ImageNet).
  • CNNs agree with each other even more, across architectures: CNN–CNN κ about 0.45–0.67 by experiment. The highest pair is κ = 0.793 (DenseNet-121 against ResNet-18, different families and depths). This holds although humans were more accurate: silhouettes, human 0.75 against CNN 0.54.
  • CNN–human consistency is near chance for cue-conflict and edge stimuli.
    • It does not improve with ImageNet accuracy (R² = 0.001 and 0.003). For silhouettes it rises with accuracy (R² = 0.253); for ImageNet images it falls (R² = 0.214).
    • "AlexNet from 2012 is just as error-consistent as recent models."
  • The recurrent "current best model of the primate ventral stream" behaves like a feedforward ResNet-50.
    • Cue conflict: CORnet-S–human κ = .066, ResNet-50–human .068, AlexNet .080, human–human .331.
    • CORnet-S against ResNet-50 κ = .711.
    • Recurrence alone did not make behaviour human-like.
  • Models with more shape bias (trained on stylised images) were more error-consistent with humans on cue-conflict stimuli.
  • The authors' principles:
    • "Making consistent errors on a trial-by-trial basis is a necessary condition" for similar strategies, though not a sufficient one.
    • "Human-level accuracy does not imply human-like decision making."
    • Errors are the informative trials when both systems are near ceiling, hence the name.
    • A metric a model was optimised for (here Brain-Score) can look good while an unoptimised behavioural metric exposes the gap.

Limits

  • The chance model treats each observer's correct and wrong trials as independent of the item (a binomial with fixed accuracy). It does not account for item difficulty. κ between two humans therefore mixes shared item difficulty with shared strategy. Both are "beyond chance" here.
  • Binary right/wrong only: which wrong answer is ignored (confusion matrices are shown separately). Small numbers of trials per condition (160) make κ noisy near ceiling, where its denominator shrinks.
  • Vision with 10 observers, rapid presentation; no repeat sessions of the same observer, so no within-person test–retest κ.

What it means for Kurisutina (inference)

  • The core fidelity metric for a capability profile is accuracy-corrected item agreement. Compare replica and person on the same items, by κ on right/wrong, and preferably on the specific wrong answer. Matching the person's score is necessary and nowhere near sufficient.
  • Calibrate against human references:
    • person against own retest: the ceiling. Not in this paper; it needs a repeat session.
    • person against other people of similar level: the population floor, about 0.3–0.5 here.
    • replica against the person must beat replica against other people (individuation) and approach the retest ceiling.
    • Because human–human κ already contains shared item difficulty, person-specific fidelity is the excess over what an item-difficulty model predicts. That is where item response theory at person level comes in (my extension, not in the paper).
  • A warning for a trained replica: networks converge on a shared machine strategy regardless of architecture (κ up to 0.79 between families), distinct from humans. A replica fine-tuned on general data will tend to share the machine pattern. Only training on the person's own item-level behaviour moves it (compare Maia, mcilroyyoung2020_maia.md).
  • Recurrence and capability settings do not by themselves produce human-like errors (CORnet-S). Error patterns must be tested directly, not assumed from architecture.
  • Use the corrected confidence intervals (simulation or the corrected formula), and shuffle item order across sessions to avoid fatigue-driven artefacts.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.