Reid McIlroy-Young (Toronto), Siddhartha Sen (Microsoft Research), Jon Kleinberg (Cornell), Ashton Anderson (Toronto). KDD 2020, doi 10.1145/3394486.3403219; arXiv:2006.01855v3 (14 July 2020). Code and models: github.com/CSSLab/maia-chess (not inspected). Read by researcher R6 (capability), 6 October 2026. Provenance: papers/capability/mcilroyyoung2020_maia.provenance.json.
What was read
- Read in full: every line of the pdftotext conversion (1,888 lines, 11 pages): sections 1–6, references, and the supplement (7.1 move-prediction training and architecture; 7.2 conversion from centipawns to win probability; 7.3 individual blunder models; 7.4 grouped blunder models; data availability). Tables 1–5 included.
- Not read: Figures 1–5, 7, 8 and 11 as images; their key values are stated in the text and Table 1. Figure 6 is the per-move agreement matrix; its values were read from the text layer, with the extremes confirmed by the text (65–79% between Maia models).
What they did
- The task is to predict the move a human plays, not the best move, "at a specific skill level, in a tuneable way".
- Data: Lichess games, rated with Glicko-2 (mean rating 1525; 5% below 1000, 10% above 2000).
- Bullet games and moves made with less than 30 s left are excluded; the first 10 ply are discarded.
- Nine test sets, one per 100-point bin from 1100 to 1900, each 10,000 games (about 500,000 positions) where both players are in the bin.
- Baselines, each a one-dimensional way to weaken a strong engine:
- Stockfish limited to search depths 1–15, the standard way of making weaker sparring engines;
- Leela (an open AlphaZero) at 8 checkpoints of different strength.
- Maia: the AlphaZero/Leela network, trained on human games instead of self-play.
- No tree search at prediction time.
- Input includes the previous 12 ply.
- One model per rating bin, each on 12 million games.
- Size: 6 residual blocks × 64 channels. A 24×320 network gave "a small performance boost at a significant cost in compute".
- Second task: predict whether the next move is a blunder (it loses ≥ 10 points of win probability), from the board alone or with metadata such as rating and clock. Also: whether more than 10% of people who reached a common position blunder there.
Main results (verified)
- Weakening an engine does not make it human.
- Depth-limited Stockfish matched 33–41% of human moves.
- Every depth's accuracy rises with the human's rating. Depth 15 matches 1900-rated players 5 points more often than 1100-rated ones.
- "If your goal is to use Stockfish to predict the moves of even a relatively weak human player, you will achieve the best performance by choosing the strongest version."
- Depths 11, 13 and 15 perform almost identically despite large strength differences: "as Stockfish increases (or decreases) in strength, it does so largely orthogonally from how humans increase or decrease in strength".
- Early checkpoints are not human either. Leela peaked at 46%, with flat curves (Leela 2700 matched 40% at every level from 1100 to 1900).
- Training on a level's own games works.
- Maia's accuracy ranged from 46% (Maia 1900 predicting 1100s) to 52.9%, above every Stockfish and Leela model throughout 1100–1900.
- Each model's curve is unimodal, peaking at or one bin above its training level (Table 1): Maia 1100 peaks at 1200 (50.8%), 1300 at 1400 (51.8%), 1500 at 1700 (52.2%), 1700 at 1800 (52.7%), 1900 at 1900 (52.9%).
- For comparison, Stockfish models peak at 1900 at every depth (35.4–41.1%); Leela models peak at an end of the range (26.3–46.0%).
- Search hurts human prediction.
- Ten rollouts of tree search cut accuracy by 5–10 points against using the network alone. More rollouts did not change it.
- The 12-ply history added about 3 points.
- Maia models agree with each other 65–79% of the time, more than with any non-Maia model. Human moves at all levels occupy "a distinct subspace". Stockfish and Leela versions agree with each other only when similar in strength.
- Maia predicts human errors. Maia beats every Leela model across the range of move quality, including blunders. A dedicated residual CNN predicted individual blunders at 67.7% (board only) and 71.7% (with metadata), against 56.4% and 63.0% for random forests. For popular positions, whether more than 10% of players blunder: 76.9%.
- The authors' framing: "approximating human performance should not simply mean matching overall performance"; we need systems that "emulate human performance on an instance-by-instance basis". They suggest finer alignment "may begin identifying novel distinctions among human players with the same rating".
Limits
- One domain (online blitz and rapid chess), a 1100–1900 band, and one-dimensional skill (rating).
- Accuracy is top-1 move matching. Calibration of the predicted distribution is not reported.
- About 53% may be near the ceiling given human randomness. No human test–retest or between-human agreement is given as a reference.
- Bins model groups, not individuals; the follow-up paper does individuals (
mcilroyyoung2022_individual_chess.md).
What it means for Kurisutina (inference)
- This is the central evidence that "settings lower skill faithfully" is wrong as stated.
- A strong system turned down along its natural dials (search depth, training time) reaches a weaker score, but does not make a weaker person's moves. Its errors are not theirs.
- Human level was reproduced by training on that level's behaviour: content, not a capacity dial.
- For Amadeus: a person's level and error pattern must come from their data (slot content). A capacity or loop-count setting may be needed, but it cannot be the main mechanism for matching a person.
- More thinking can make a replica less faithful. Search raised strength and lowered human-likeness, because blitz players do not search that deeply. A replica that "thinks longer than the person" overshoots. Its deliberation must be matched to the person's time and process.
- A small network suffices for a skill level. Maia's network (6×64 residual blocks, roughly 10⁶ parameters by my arithmetic, not stated) captured a rating level's move distribution, and bigger networks added little. At 12M games per level, the bottleneck is data and objective, not size.
- How to measure fidelity. Report move-matching as a curve across skill levels, and require unimodality at the person's level: the replica should match its own level best, and others less. Also measure prediction of the person's errors (blunders) and agreement matrices between replicas. Aggregate strength (rating) alone is insufficient.