Brooke N. Macnamara (Princeton), David Z. Hambrick (Michigan State), Frederick L. Oswald (Rice). Psychological Science 25(8):1608–1618, doi 10.1177/0956797614535810, with the 2018 corrigendum (doi 10.1177/0956797618769891). Read as the file hosted by a Purdue skill-learning lab: the corrigendum prepended to the article, with corrected values overlaid on the original text. Read by researcher R6 (capability), 6 October 2026. Provenance: papers/capability/macnamara2014_practice_meta.provenance.json.
What was read
- Read in full: the corrigendum (3 pages, Table 1 of corrected results) and the article (11 pages): introduction, method, results, moderator and additional models, publication bias, discussion, notes and references, plus the Figure 1–3 text layers.
- Not read: the supplemental material (Table S1 reliability scenarios, Figures S1–S16) and the open data (osf.io/rhfsk).
- Values below are the corrected ones unless marked "original". The pdftotext layer interleaves old and new numbers, so the corrigendum's own table was used to tell them apart.
What they did
- Question: how much of the variance in performance does accumulated deliberate practice explain? The comparison is with Ericsson et al.'s (1993) claim that it "largely" accounts for individual differences.
- Studies: 9,331 records searched up to March 2014 → 88 studies, 111 independent samples, 157 effect sizes, N = 11,135.
- Domains: music, games, sports, education, professions.
- Effect size: the correlation between hours of practice and performance (Cohen's d converted to biserial r).
- Random-effects model, with moderators: domain, task predictability, how practice and performance were measured.
- Publication bias checked by funnel plot and trim-and-fill.
Main results (verified; corrected values)
- Overall: r = .38 [.33, .42]. Practice explains 14% of the variance in performance [11%, 18%] (original: r = .35, 12%). I² = 88.5, so heterogeneity is high.
- By domain (variance explained):
- games 24% (r = .49);
- music 23% (r = .48);
- sports 20% (r = .45);
- education 5% (r = .22);
- professions 1% (r = .09, n.s.).
- Originally reported: 26%, 21%, 18%, 4% and < 1%.
- By task predictability: high 24%, moderate 14%, low 6%.
- By how practice was measured: retrospective interview 20%, questionnaire 15%, log (recorded as it happened) 4%. "This finding suggests that … a 'high-fidelity' approach … might reveal that the relationship … is weaker."
- By how performance was measured: group membership 26%, laboratory tasks 12%, expert ratings 11%, standardised objective scores such as chess ratings 10%.
- Correcting for unreliability (original analysis, assuming .80 reliability for both measures): r = .43, so 19% of reliable variance. That is games 41%, music 33%, sports 28%, education 7%, professions < 1%.
- Robust to model choice. Excluding team sports, or keeping only solitary practice, gives 11–14% overall. Solitary practice is not a stronger predictor.
- No publication bias detected.
- Other factors the authors cite (not meta-analysed here):
- Master-level chess players' practice ranged from about 3,000 to more than 23,000 hours (Gobet & Campitelli 2007).
- Practice accounted for about a third of the reliable variance in chess and music (Hambrick et al. 2014).
- Starting age predicts chess rating after controlling for practice.
- General intelligence predicts performance across domains.
- Working memory capacity predicted piano sight-reading beyond deliberate practice, with no interaction. It was "as important a predictor … for beginning pianists as it was for pianists who had engaged in thousands of hours" (Meinz & Hambrick 2010).
Limits
- The data are correlational and between persons. A within-person practice effect could differ.
- Retrospective estimates dominate (8,233 of 11,135 participants answered questionnaires). The less biased log measure gives the smallest effect.
- "Deliberate practice" is defined differently across studies, and is least well defined in education and professions.
- Range restriction among elite samples is not modelled. The authors' account of the unexplained variance (abilities, starting age) relies on cited single studies, not on this meta-analysis.
What it means for Kurisutina (inference)
- A person's skill level cannot be inferred from their practice history. Even in games, hours explain about a quarter of the variance (about 40% of the reliable variance). Two people with the same history can sit far apart. A slot must be calibrated on the person's measured performance, not built from "how much they practised".
- General capacities matter beyond content. The working-memory result (cited, not meta-analysed) says a domain-general capacity keeps predicting performance after thousands of hours. This supports part of the working answer: some capacity really is a person-level setting. A replica of a high-capacity expert needs a base whose capacity range reaches that person.
- But content dominates the expert–novice gap, in the sense that the domain's practised knowledge (chunks, templates; see
gobet2000_five_seconds_or_sixty.md) is what practice builds. - The two layers interact. People of equal capacity diverge through what they learned, and people with equal practice diverge through capacity. A design that puts only content in the slot, or only capacity in settings, leaves part of each person's level unexplained.