Kurisutina

Aligning superhuman AI with human behavior: chess as a model system ("Maia")

Reid McIlroy-Young (Toronto), Siddhartha Sen (Microsoft Research), Jon Kleinberg (Cornell), Ashton Anderson (Toronto). KDD 2020, doi 10.1145/3394486.3403219; arXiv:2006.01855v3 (14 July 2020). Code and models: github.com/CSSLab/maia-chess (not inspected). Read by researcher R6 (capability), 6 October 2026. Provenance: papers/capability/mcilroyyoung2020_maia.provenance.json.

What was read

  • Read in full: every line of the pdftotext conversion (1,888 lines, 11 pages): sections 1–6, references, and the supplement (7.1 move-prediction training and architecture; 7.2 conversion from centipawns to win probability; 7.3 individual blunder models; 7.4 grouped blunder models; data availability). Tables 1–5 included.
  • Not read: Figures 1–5, 7, 8 and 11 as images; their key values are stated in the text and Table 1. Figure 6 is the per-move agreement matrix; its values were read from the text layer, with the extremes confirmed by the text (65–79% between Maia models).

What they did

  • The task is to predict the move a human plays, not the best move, "at a specific skill level, in a tuneable way".
  • Data: Lichess games, rated with Glicko-2 (mean rating 1525; 5% below 1000, 10% above 2000).
    • Bullet games and moves made with less than 30 s left are excluded; the first 10 ply are discarded.
    • Nine test sets, one per 100-point bin from 1100 to 1900, each 10,000 games (about 500,000 positions) where both players are in the bin.
  • Baselines, each a one-dimensional way to weaken a strong engine:
    • Stockfish limited to search depths 1–15, the standard way of making weaker sparring engines;
    • Leela (an open AlphaZero) at 8 checkpoints of different strength.
  • Maia: the AlphaZero/Leela network, trained on human games instead of self-play.
    • No tree search at prediction time.
    • Input includes the previous 12 ply.
    • One model per rating bin, each on 12 million games.
    • Size: 6 residual blocks × 64 channels. A 24×320 network gave "a small performance boost at a significant cost in compute".
  • Second task: predict whether the next move is a blunder (it loses ≥ 10 points of win probability), from the board alone or with metadata such as rating and clock. Also: whether more than 10% of people who reached a common position blunder there.

Main results (verified)

  • Weakening an engine does not make it human.
    • Depth-limited Stockfish matched 33–41% of human moves.
    • Every depth's accuracy rises with the human's rating. Depth 15 matches 1900-rated players 5 points more often than 1100-rated ones.
    • "If your goal is to use Stockfish to predict the moves of even a relatively weak human player, you will achieve the best performance by choosing the strongest version."
    • Depths 11, 13 and 15 perform almost identically despite large strength differences: "as Stockfish increases (or decreases) in strength, it does so largely orthogonally from how humans increase or decrease in strength".
  • Early checkpoints are not human either. Leela peaked at 46%, with flat curves (Leela 2700 matched 40% at every level from 1100 to 1900).
  • Training on a level's own games works.
    • Maia's accuracy ranged from 46% (Maia 1900 predicting 1100s) to 52.9%, above every Stockfish and Leela model throughout 1100–1900.
    • Each model's curve is unimodal, peaking at or one bin above its training level (Table 1): Maia 1100 peaks at 1200 (50.8%), 1300 at 1400 (51.8%), 1500 at 1700 (52.2%), 1700 at 1800 (52.7%), 1900 at 1900 (52.9%).
    • For comparison, Stockfish models peak at 1900 at every depth (35.4–41.1%); Leela models peak at an end of the range (26.3–46.0%).
  • Search hurts human prediction.
    • Ten rollouts of tree search cut accuracy by 5–10 points against using the network alone. More rollouts did not change it.
    • The 12-ply history added about 3 points.
  • Maia models agree with each other 65–79% of the time, more than with any non-Maia model. Human moves at all levels occupy "a distinct subspace". Stockfish and Leela versions agree with each other only when similar in strength.
  • Maia predicts human errors. Maia beats every Leela model across the range of move quality, including blunders. A dedicated residual CNN predicted individual blunders at 67.7% (board only) and 71.7% (with metadata), against 56.4% and 63.0% for random forests. For popular positions, whether more than 10% of players blunder: 76.9%.
  • The authors' framing: "approximating human performance should not simply mean matching overall performance"; we need systems that "emulate human performance on an instance-by-instance basis". They suggest finer alignment "may begin identifying novel distinctions among human players with the same rating".

Limits

  • One domain (online blitz and rapid chess), a 1100–1900 band, and one-dimensional skill (rating).
  • Accuracy is top-1 move matching. Calibration of the predicted distribution is not reported.
  • About 53% may be near the ceiling given human randomness. No human test–retest or between-human agreement is given as a reference.
  • Bins model groups, not individuals; the follow-up paper does individuals (mcilroyyoung2022_individual_chess.md).

What it means for Kurisutina (inference)

  • This is the central evidence that "settings lower skill faithfully" is wrong as stated.
    • A strong system turned down along its natural dials (search depth, training time) reaches a weaker score, but does not make a weaker person's moves. Its errors are not theirs.
    • Human level was reproduced by training on that level's behaviour: content, not a capacity dial.
    • For Amadeus: a person's level and error pattern must come from their data (slot content). A capacity or loop-count setting may be needed, but it cannot be the main mechanism for matching a person.
  • More thinking can make a replica less faithful. Search raised strength and lowered human-likeness, because blitz players do not search that deeply. A replica that "thinks longer than the person" overshoots. Its deliberation must be matched to the person's time and process.
  • A small network suffices for a skill level. Maia's network (6×64 residual blocks, roughly 10⁶ parameters by my arithmetic, not stated) captured a rating level's move distribution, and bigger networks added little. At 12M games per level, the bottleneck is data and objective, not size.
  • How to measure fidelity. Report move-matching as a curve across skill levels, and require unimodality at the person's level: the replica should match its own level best, and others less. Also measure prediction of the person's errors (blunders) and agreement matrices between replicas. Aggregate strength (rating) alone is insufficient.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.