Google DeepMind and Google. NeurIPS 2024; arXiv:2402.04494v2 (21 October 2024). Dataset, weights and code: github.com/google-deepmind/searchless_chess (not inspected). Read by researcher R6 (capability), 6 October 2026. Provenance: papers/capability/ruoss2024_grandmaster_chess.provenance.json.
What was read
- Read in full: every line of the pdftotext conversion (2,442 lines, 26 pages): sections 1–7, references, Appendices A (setup, dataset statistics, Stockfish, AlphaZero and Leela setups, compute) and B (additional ablations, inference times, legal moves, loss curves, prediction targets, Fischer random chess, tactics, playing style with game scores), and the NeurIPS checklist.
- Not read: Figures 2–4 and A1–A7 as images; values used are in the tables and text.
What they did
- ChessBench: 10M Lichess games (February 2023). Every board was annotated by Stockfish 16 at 50 ms per legal move:
- 530M board states and 15.3B action-values, about 8,864 days of single-core Stockfish time;
- the "oracle" plays at about Lichess blitz Elo 2713.
- Test sets: 1K games from a different month (14.7% of boards also appear in training, mostly openings) and 10K rated puzzles (1.33% overlap).
- Models: decoder-only transformers of 9M (8 layers, width 256), 136M (8 layers, width 1024) and 270M (16 layers, width 1024).
- Trained by supervised learning to predict binned action-values (HL-Gauss loss, 128 bins) for 2.67 epochs, on 128 TPU v5 chips per large model.
- The policy picks the move with the highest predicted value. No search at test time.
- Compared with:
- Stockfish 16;
- AlphaZero (27.6M parameters, 44M games) and Leela Chess Zero (T82), each with and without 400-simulation tree search;
- GPT-3.5-turbo-instruct.
- Ablations on the 9M model: prediction target (action-value, state-value, behavioural cloning), loss, depth at fixed parameters, number of bins, data sampler, architecture (ConvNet).
Main results (verified)
Playing strength (Table 1):
| Agent | Search | Tournament Elo | Lichess vs bots | Lichess vs humans | Puzzles |
|---|---|---|---|---|---|
| 9M transformer | no | 2025 | 2054 | – | 88.9% |
| 136M transformer | no | 2259 | 2156 | – | 94.5% |
| 270M transformer | no | 2299 | 2299 | 2895 | 95.4% |
| AlphaZero policy net | no | 1777 | – | – | 56.1% |
| AlphaZero, 400 MCTS | yes | 2470 | – | – | 95.6% |
| Leela policy net | no | 2292 | 2224 | – | 88.6% |
| Leela, 400 MCTS | yes | 2858 | 2620 | – | 99.6% |
| Stockfish 16, 50 ms per move (oracle) | yes | 2711 | 2713 | – | 99.8% |
| Stockfish 16, 1.5 s per board | yes | 2935 | 2940 | – | 100% |
| GPT-3.5-turbo-instruct | – | – | 1755 | – | 66.5% |
- Grandmaster-level blitz against humans (2895, from 174 games), with no search. Against bots the same model rates 2299: "most losses against bots can be explained by just one tactical blunder".
- Search is worth hundreds of Elo for policy networks. AlphaZero gains 693 tournament points from 400 simulations; Leela 566. The searchless transformer nearly matches AlphaZero with search on puzzles, but "perfect distillation is still beyond reach".
- Scale and data both matter. Bigger models and bigger data sets do better. On 10K games, models of 7M parameters and above overfit; not on 100K or 1M games.
- Depth at fixed parameters: puzzle accuracy 54.7 (2 layers), 40.5 (4), 79.5 (8), 81.3 (16), 79.5 (32). "Depth is important, but not beyond a certain point" (saturates at about 16).
- Behavioural cloning learns less. Trained only on the oracle's best move, it lags state-value training on the same data (puzzles 65.7% against 77.5% at 9M), because it discards the information in the values.
- Generalisation is bounded by the training distribution. The 9M model plays Fischer random chess at 1539 Elo, against 2054 in standard chess.
- Style: masters judged it aggressive and "optimistic", favouring moves that give opponents difficult decisions. It is "more enjoyable than playing a normal engine". It disagrees with Stockfish in telling places.
- Inference cost: 57 ms per move for 270M on a TPU v5, against about 1.5 s for Stockfish's 50-ms-per-move oracle.
Limits
- The teacher is a superhuman search engine with 15B labels. This shows how much skill a network can absorb when supervised move by move; it does not show learning from sparse human data.
- The Elo against humans is from 174 games. It was measured only for the largest model; the other models were rated against bots.
- No history input (FEN only): threefold-repetition and "indecisiveness" workarounds were needed, including deferring to Stockfish in won positions.
- Comparisons across engines mix inputs (FEN against PGN), training (supervised learning against RL) and search.
What it means for Kurisutina (inference)
- Expert-level skill in one domain fits in 10⁷–10⁸ parameters.
- 9M parameters reach about 2000–2050 (strong club level); 270M reach grandmaster blitz against humans.
- These networks hold only chess. At Allen-Zhu's 2 bits per parameter (
summaries/base/allenzhu2024_knowledge_capacity.md), 9M parameters could hold up to about 1.8×10⁷ bits. That is the order of a master's chess pattern store estimated from Gobet & Simon (gobet2000_five_seconds_or_sixty.md, my arithmetic there). - This is why a large general model is not needed to match a human expert in one domain. It needs domain-dense training and enough depth.
- Search (thinking time) substitutes for network quality, partially.
- Hundreds of Elo for weak policies; the searchless 270M approaches searched AlphaZero.
- This complements Maia: search raises strength but lowers human-likeness. A replica must match the person's search, not maximise it.
- The skill is content-bound in networks as in people. Out-of-distribution play (Fischer random) collapses, as experts' recall advantage shrinks on random positions. Domain skill is learned patterns, not general capacity, in both.
- Depth is a capacity parameter that saturates. It is worth sizing per domain, about 16 layers here. It is a base-level property, not slot content.