Kurisutina

Reproducible brain-wide association studies require thousands of individuals

What was read

  • Full text from PubMed Central (PMC8991999): abstract, main text, all Methods sections, author notes. Read in full.
  • NOT read: Extended Data figure legends and Supplementary Information (not in the PMC XML). A correction was published 9 May 2022 (10.1038/s41586-022-04692-3); its content is not verified here.

Question

How large are the associations between inter-individual differences in MRI measures (cortical thickness, resting-state functional connectivity, task activation) and complex phenotypes (cognitive ability, psychopathology), and how big must a sample be for such an association to replicate? The paper defines these as brain-wide association studies (BWAS): "studies of the associations between common inter-individual variability in human brain structure/function and cognition or psychiatric symptomatology". The median neuroimaging study has about 25 participants.

Data

  • ABCD (n = 11,874 children aged 9–10, 21 sites); after strict motion denoising (≥ 8 min at filtered framewise displacement < 0.08 mm) n = 3,928, split into matched discovery and replication halves of 1,964.
  • HCP (n = 1,200 adults 22–35, single scanner, 60 min of rest per person), n = 900 used.
  • UK Biobank (n = 35,735 adults 40–69, 6 min of rest), n = 32,572 used.
  • 41 phenotypes (NIH Toolbox cognition, CBCL psychopathology, demographics). Brain features: cortical thickness at 59,412 vertices, 333 ROIs, 13 networks; RSFC at 77,421 edges, 105 network pairs, 100 principal components. Around 11 million univariate associations; bootstrapped subsamples from n = 25 to the full sample, 1,000 resamples per size (100 for multivariate).

Results (with numbers)

  1. Effects are tiny. Across all univariate brain–phenotype associations in ABCD, the median |r| is 0.01. The top 1% of all associations exceed |r| = 0.06. The largest univariate association that replicated out of sample is |r| = 0.16. Adjusting for sociodemographic covariates lowers the strongest effects further (top 1% Δr = −0.014).
  2. Sampling variability swamps them at typical sizes. At n = 25 the 99% confidence interval of a univariate association is r ± 0.52; two n = 25 subsamples can reach opposite conclusions about the same association. Even in split halves of n = 1,964, the top 1% of effects remain inflated by r = 0.07 (78%) on average.
  3. The effect-size distribution replicates across datasets, ages, sites and scanners: at n = 900 the top 1% threshold is |r| > 0.11 (ABCD), 0.12 (HCP), 0.09 (UKB), with median |r| = 0.02–0.03 for fluid intelligence vs RSFC. Task-fMRI activation contrasts (86 contrasts × 39 phenotypes in HCP) have the same distribution as RSFC. The authors call the distribution "universal to BWAS with current technologies and methods".
  4. Reliability is not the main limit. Within-person reliability: NIH Toolbox r = 0.90, CBCL 0.94, cortical thickness > 0.96, RSFC 0.48 (ABCD), 0.79 (HCP), 0.39 (UKB). Only RSFC has headroom, and "theoretical maximum BWAS effect sizes are unlikely to be reached owing to fundamental biological limits on the strength of the true association and/or the limitations of behavioural phenotyping and MRI physics".
  5. Statistical error rates. At n = 1,000 the false-negative rate is 75–100% and half of significant associations are inflated by at least 100%. Maximum power at n = 3,928 is 0.68. Replication (same sign, both significant at P < 0.05) succeeds 25% of the time at n = 1,964 and around 5% at n < 500. Bonferroni correction (P < 10⁻⁷) needs n = 9,500 for 80% power on the top 1% of effects, versus n = 2,200 uncorrected.
  6. The underpowered-BWAS paradox: at small n only inflated effects reach significance, stricter thresholds select for more inflated effects, and regression to the mean makes replication failure the most likely outcome. Publication bias then feeds inflated estimates into later power analyses.
  7. Multivariate is better but still needs thousands. SVR and CCA on PCA-reduced features: out-of-sample associations average r = 0.17 versus in-sample 0.46 (a 63% drop). Best case: RSFC predicting crystallized intelligence, out-of-sample rpred = 0.39 at n ≈ 2,000 (univariate max 0.16). RSFC beats cortical thickness; cognitive tests beat questionnaires. Out-of-sample multivariate strength tracks the univariate effect size (r = 0.79 across phenotypes). Low-dimensional feature spaces replicate best, meaning the associations are distributed widely rather than localised.
  8. Small samples remain appropriate for non-BWAS designs: group-average functional organisation is stable at n = 25; precise individual-specific maps come from repeatedly scanning the same person; lesion, task, longitudinal and interventional (within-person) designs have larger effects and higher reliability.

Recommendations: BWAS need thousands of standardly processed participants; report in-sample and out-of-sample effect sizes; prefer within-person and interventional designs; prefer function over structure and direct behavioural measures over questionnaires.

Limits

  • Cross-sectional, mostly linear methods; the paper measures the association distribution "with current technologies and methods", not a physical ceiling. The authors themselves list both biological limits and measurement limits as possible reasons.
  • RSFC reliability in the largest dataset (UKB, 6 min) is 0.39, so UKB effects are attenuated by measurement more than the others.
  • Child cohort for the main analyses; adult cohorts used for verification of the distribution, not the full error analysis.
  • Extended Data and Supplementary figures, including the analytic power estimates and the correction notice, were not read.

What the brief uses it for, and whether it holds

Brief section 2: "at MRI resolution, human brain–trait associations are weak enough that reproducible studies need thousands of subjects [3]". Brief section 11: "Trait ceilings are method ceilings, not physical bounds. Stated as such (section 2)."

  • The first claim holds exactly. Median |r| = 0.01, largest replicating |r| = 0.16, replication rates near 5% below n = 500, and the paper's own conclusion is that thousands are required.
  • The second claim is half-supported. Marek et al. say the ceilings are "unlikely to be reached owing to fundamental biological limits ... and/or the limitations of behavioural phenotyping and MRI physics". They do not decide between the two. The brief's "method ceilings, not physical bounds" is a permissible reading, not a finding.
  • For Variant A (section 8, "trait constraints at the resolution of cohort labels"): structural MRI is the weakest modality here, and questionnaire-type mental-health phenotypes are the weakest outcomes. Whatever a post-mortem structural measurement could supply about a person, this paper says the cross-sectional map from structure to trait is fitted from effects of order r ≈ 0.1 and needs thousands of paired brains to estimate. A brain-bank effect-size study would be fighting exactly these numbers.
  • For Variant B, the paper is quietly on the brief's side: the designs it endorses (within-person, repeated sampling of one individual, interventions) are the designs the brief proposes (person-specific decoders, longitudinal enrolment, controlled experiences with predicted change). Marek's point that precise individual maps come from many hours on one person is the same lesson as Tang's person-specific decoders [4].

Cross-references

  • [2] Linneweber 2020: the fly result (r ≈ −0.67 at N = 103) is what a structure-to-trait association looks like when the right circuit is measured directly; Marek is what it looks like at MRI resolution in humans.
  • [4] Tang 2023: the within-person decoding route that Marek's small-sample section implicitly endorses.
  • [12] Park 2026 and [13] Anderson 2025: both report person-level prediction from within-person or cross-person data; Marek's inflation numbers are a warning for any small-n claim in the brief's own pilots (two or three participants in step 1, "a handful" in step 3).

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.