Kurisutina

BrainLM: a foundation model for brain activity recordings

ICLR 2024 conference paper (OpenReview proceedings PDF, 12 pages). Read by the main session for research direction R3, 4 October 2026. Provenance: papers/amadeus/ortegacaro2024_brainlm.provenance.json.

What was read

  • Read in full: the 12-page main paper, from pdftotext -layout: abstract, introduction, related work, methods, results, discussion, references; tables 1–3 and figure captions.
  • Not read: the supplementary material. It holds the implementation details, Table 14 (latent-space clinical information), Supplementary Table 17 (zero-shot regression by model size) and Supplementary Figure 8 (attention for other variables), all referred to in the text. The figures were not seen as images.

What they did

  • Data. UK Biobank, 76,296 task and resting fMRI recordings (ages 40–69, TR 0.735 s), and the Human Connectome Project, 1,002 recordings (TR 0.72 s). Each recording was reduced to 424 parcels (AAL-424 atlas) at about 1 Hz and robust-scaled per parcel.
  • Training split. "Trained on 80% of the UKB dataset (61,038 recordings)", evaluated on the held-out 20% and on HCP.
  • Model. A masked autoencoder Transformer, BERT and MAE style:
    • 200-step windows are split into patches of 20 time steps; 20%, 75% or 90% of patches are masked;
    • a 4-layer encoder (4 heads) and a 2-layer decoder reconstruct the masked signal under MSE loss.
    • Sizes: 13M, 111M and 650M parameters.
  • Downstream tasks:
    • fine-tune with an MLP head to regress age, neuroticism, PTSD (PCL-5) and anxiety (GAD-7);
    • fine-tune to forecast the next 20 time steps from 180;
    • attention analysis;
    • k-NN classification of parcels into 7 functional networks.

Main results (verified)

  • Masked reconstruction: R² 0.464 on held-out UK Biobank, 0.278 on HCP. Scaling with data and model size is shown in Fig. 4, as a figure only.
  • Forecasting 20 s ahead is weak. UK Biobank R² 0.098 (650M), 0.095 (111M) and 0.086 (13M); HCP 0.061, 0.056 and 0.028. The same Transformer without pretraining scored 0.012; the LSTM and Latent ODE about 0. Pretraining helps clearly, but under a tenth of the variance of the next 20 s of parcel activity is predicted.
  • Clinical regression (MSE, lower is better). BrainLM beat the baselines on all four variables.
    • Age, z-scored: 0.464 (13M) and 0.503 (111M), against 0.596 for the LSTM and 0.659 for the SVR.
    • PTSD: 0.018 and 0.015, against 0.019–0.022.
    • Anxiety: 0.074 and 0.073, against 0.081–0.090.
    • Neuroticism: 0.072 and 0.069, against 0.076–0.087.
    • The gains are modest. The larger model is worse on age.
  • Functional networks: 58.8% accuracy over 7 networks with k-NN on attention weights, against 49.4% (VAE), 39.2% (raw data) and 25.9% (GCN).

Limits and problems I found

  • Possible train/test overlap with HCP. The paper says the model was trained on 80% of UK Biobank and evaluated on "the full HCP dataset". It also says "our training dataset comprised 6,700 hours … from 77,298 recordings across the two repositories", and 77,298 = 76,296 + 1,002, all of both. If HCP was in pretraining, the "never-before-seen cohort" generalisation claim (R² 0.278) is not clean. The main text does not resolve it.
  • "Zero-shot" network identification is supervised at the classifier. Pretraining had no network labels, but the k-NN classifier was trained on 80% of the parcels' network labels in each recording.
  • Mixed normalisations in Table 2: "MSEs for the larger BrainLM models are not comparable to other models".
  • Two tokenisation schemes are described (per-parcel 20-step patches; multi-parcel 2D patches by y-coordinate) without saying which model used which.
  • Weak baselines. The "raw data" baseline for age has MSE 2.0 on a z-scored target, worse than predicting the mean (about 1.0). No R² or correlation is reported for the clinical variables.
  • A population model. Nothing is person-specific. Individual differences enter only as regression targets.
  • fMRI itself: a slow haemodynamic signal (about 1 Hz here) over 424 parcels.

What it means for Amadeus (inference)

  • A "base" for human brain dynamics exists, and it is weak at what a model of one brain most needs. Self-supervised pretraining on thousands of hours helps (forecast R² 0.09 against 0.01 without pretraining), but the next 20 seconds of whole-brain parcel activity are mostly unpredicted. Either fMRI at this resolution is noise-limited, or the dynamics depend on things the signal does not carry: the content of thought, external input. Both point the same way. Coarse fMRI dynamics alone will not carry a person's thinking; anchors that tie activity to content are needed (Goldstein 2022, MindEye2).
  • Scale comparison for the month idea. BrainLM's whole corpus is 6,700 hours across tens of thousands of people. A month of continuous recording of one person is about 720 hours, roughly a tenth of that corpus, all from one brain. This is not possible with fMRI: a person cannot stay in a scanner for a month (R1 covers what dense sampling has achieved). With wearable or implanted recording, one person's month would be a large dataset by current standards.
  • Self-supervised then fine-tuned is the right recipe for the "model of one brain" that the architecture doc proposes: pretrain on the person's own unlabeled stream, then tie to meaning through labelled anchors. The evidence here is only that pretraining helps downstream tasks; nothing shows it captures a person.
  • Caution on claims: the paper's framing ("interpretable", "zero-shot", "never-before-seen") is stronger than its evidence. Its numbers should be cited, not its adjectives.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.