Reid McIlroy-Young, Russell Wang (Toronto), Siddhartha Sen (Microsoft Research), Jon Kleinberg (Cornell), Ashton Anderson (Toronto). KDD 2022, doi 10.1145/3534678.3539367; arXiv:2008.10086v3 (15 June 2022). Code: github.com/CSSLab/maia-individual (not inspected). Read by researcher R6 (capability), 6 October 2026. Provenance: papers/capability/mcilroyyoung2022_individual_chess.provenance.json.
What was read
- Read in full: every line of the pdftotext conversion (1,778 lines, 12 pages): sections 1–6 including ethics, references, and the supplement (dataset criteria, Tables 4–7, gradient-depth and learning-rate tuning, starting-model choice, calibration, stylometry by ply, the proof of the Chernoff bound).
- Not read: Figures 1–11 as images; values used are those in the text and tables.
What they did
- Base: Maia 1900, the population model trained on 12M games of 1900-rated players (
mcilroyyoung2020_maia.md). - Personalisation ("Transfer Maia"): fine-tune the whole network on one player's blitz games.
- Learning rate 1e-4, 30,000 steps. Most of the gain comes in the first 12,000 steps, about 3M board–move pairs.
- Train / validation / test split 80/10/10, plus a "future" set: the player's December 2020 games, played after all training games.
- Players: Lichess blitz players rated 1000–2000 with low rating variance; titled players and bots excluded. 400 evaluation players, median rating 1739, grouped by games played (1,000 to 40,000).
- Design choices tested on exploration players:
- freezing layers only hurt;
- random initialisation hurt;
- which Maia level they started from barely mattered ("the fine-tuning process dominates this choice").
- Tasks:
- predict each move (after ply 10);
- stylometry: identify which of 400 players produced a set of games, by picking the personalised model with the most correct predictions. This was never a training objective.
Main results (verified)
-
Personalisation adds about 4 points of move matching. Over the nearest population (Maia) model, by player:
- main effect +4.186 points (SE ± 0.101);
- about the same as the gap between Maia and the strongest classical engines;
- perplexity 1.95 against 2.15 for Maia 1900 (openings excluded).
-
It needs thousands of the person's decisions (Table 2: nearest Maia → personalised):
Player's games Nearest Maia Personalised 1,000 0.527 0.497 (worse: overfitting) 5,000 0.532 0.550 10,000 0.528 0.556 20,000 0.524 0.564 30,000 0.520 0.567 40,000 0.528 0.580 - "Significant uplift … for players who have played at least 5,000 games."
- Pooling across players for small data is left as future work.
-
Better on every kind of move, including blunders. Personalised models assign more probability than the population model to the person's optimal moves, small errors and large mistakes alike.
-
Generalisation:
- Accuracy holds in positions never seen in training; by ply 30, only 0.02% of positions were seen.
- It holds on later games: the future set was 0.45–1.1 points higher than the test set.
-
Stylometry: near-unique identification. Top-1 accuracy among 400 players (Table 3):
Games used All moves Ply 10+ Ply 30+ 10 .86 .47 .11 30 .94 .81 .26 100 .98 .95 .55 - Chance is 0.25%. A Naive Bayes baseline on centipawn-loss vectors reached 1.4%.
- Using only blunders, accuracy rose to .989: "mistakes are the most discriminative moves … a player's personalized model makes blunders like that individual player."
-
Why a 4–5-point edge identifies people. By a Chernoff bound, with 400 players, 50% accuracy on others and 55% on the target, fewer than 20,000 moves suffice for 0.9 probability of correct identification.
-
Personal accuracy is nearly independent of rating in 1100–1900 (R² = 0.089).
- Preliminary personalisation of Grandmasters did not give the same gains. The authors conjecture that top players play near-optimal moves, which engines already predict.
-
Cost: about 40 minutes per player model on one V100, using about 300 MB of GPU memory.
Limits
- One domain (blitz chess) and top-1 move matching. Calibration was checked only qualitatively ("good at predicting their own accuracy").
- Stylometry is against a closed set of 400. It shows individuation, not that the model is the person.
- The population base and the person come from the same platform and time span. Long-term drift in a person is not modelled beyond one month ahead.
- Ethics, in the authors' words: deanonymisation, cheating-detection evasion, and "mimetic models" of individuals.
What it means for Kurisutina (inference)
- The closest published precedent for a skill slot. A population base plus a per-person fine-tune reproduces this person's decisions, including their characteristic mistakes, well enough to identify them among 400 from 100 games.
- That is the individuation gate (own beats other) passed decisively in a skill domain.
- The errors are the most person-specific part. A replica judged by accuracy would erase exactly what identifies the person.
- The slot needs thousands of the person's games per domain. Personalisation beat the population model only for players with at least about 5,000 games played, and was still improving at 40,000.
- Median training games in the 5,000 group: 2,432–3,399 (Table 4). That is roughly tens of thousands of moves after ply 10; the move count is my estimate.
- For a human expert, the analogue is a large, dense behavioural record in their domain, not an interview.
- With little data, use the population or level model (B8 default), as the authors also suggest.
- Personalisation overrides the starting level. Starting from Maia 1100 or 1900 made little difference. This supports content-in-the-slot, with the base supplying shared representations rather than the level.
- The top of the skill range is the hard case. At Grandmaster level, personal deviations are small relative to near-optimal play. A replica of a world-class expert needs a base strong enough to play near-optimally, plus the slot's deviations. This is where base capacity, not content, binds.
- A full fine-tune of a small base is cheap (about 40 GPU-minutes). The open question for us is adapter size: here the whole network (the 6×64 Maia) was tuned, and freezing layers hurt.