ICLR 2024 conference paper (OpenReview proceedings PDF, 12 pages). Read by the main session for research direction R3, 4 October 2026. Provenance: papers/amadeus/ortegacaro2024_brainlm.provenance.json.
What was read
- Read in full: the 12-page main paper, from
pdftotext -layout: abstract, introduction, related work, methods, results, discussion, references; tables 1–3 and figure captions. - Not read: the supplementary material. It holds the implementation details, Table 14 (latent-space clinical information), Supplementary Table 17 (zero-shot regression by model size) and Supplementary Figure 8 (attention for other variables), all referred to in the text. The figures were not seen as images.
What they did
- Data. UK Biobank, 76,296 task and resting fMRI recordings (ages 40–69, TR 0.735 s), and the Human Connectome Project, 1,002 recordings (TR 0.72 s). Each recording was reduced to 424 parcels (AAL-424 atlas) at about 1 Hz and robust-scaled per parcel.
- Training split. "Trained on 80% of the UKB dataset (61,038 recordings)", evaluated on the held-out 20% and on HCP.
- Model. A masked autoencoder Transformer, BERT and MAE style:
- 200-step windows are split into patches of 20 time steps; 20%, 75% or 90% of patches are masked;
- a 4-layer encoder (4 heads) and a 2-layer decoder reconstruct the masked signal under MSE loss.
- Sizes: 13M, 111M and 650M parameters.
- Downstream tasks:
- fine-tune with an MLP head to regress age, neuroticism, PTSD (PCL-5) and anxiety (GAD-7);
- fine-tune to forecast the next 20 time steps from 180;
- attention analysis;
- k-NN classification of parcels into 7 functional networks.
Main results (verified)
- Masked reconstruction: R² 0.464 on held-out UK Biobank, 0.278 on HCP. Scaling with data and model size is shown in Fig. 4, as a figure only.
- Forecasting 20 s ahead is weak. UK Biobank R² 0.098 (650M), 0.095 (111M) and 0.086 (13M); HCP 0.061, 0.056 and 0.028. The same Transformer without pretraining scored 0.012; the LSTM and Latent ODE about 0. Pretraining helps clearly, but under a tenth of the variance of the next 20 s of parcel activity is predicted.
- Clinical regression (MSE, lower is better). BrainLM beat the baselines on all four variables.
- Age, z-scored: 0.464 (13M) and 0.503 (111M), against 0.596 for the LSTM and 0.659 for the SVR.
- PTSD: 0.018 and 0.015, against 0.019–0.022.
- Anxiety: 0.074 and 0.073, against 0.081–0.090.
- Neuroticism: 0.072 and 0.069, against 0.076–0.087.
- The gains are modest. The larger model is worse on age.
- Functional networks: 58.8% accuracy over 7 networks with k-NN on attention weights, against 49.4% (VAE), 39.2% (raw data) and 25.9% (GCN).
Limits and problems I found
- Possible train/test overlap with HCP. The paper says the model was trained on 80% of UK Biobank and evaluated on "the full HCP dataset". It also says "our training dataset comprised 6,700 hours … from 77,298 recordings across the two repositories", and 77,298 = 76,296 + 1,002, all of both. If HCP was in pretraining, the "never-before-seen cohort" generalisation claim (R² 0.278) is not clean. The main text does not resolve it.
- "Zero-shot" network identification is supervised at the classifier. Pretraining had no network labels, but the k-NN classifier was trained on 80% of the parcels' network labels in each recording.
- Mixed normalisations in Table 2: "MSEs for the larger BrainLM models are not comparable to other models".
- Two tokenisation schemes are described (per-parcel 20-step patches; multi-parcel 2D patches by y-coordinate) without saying which model used which.
- Weak baselines. The "raw data" baseline for age has MSE 2.0 on a z-scored target, worse than predicting the mean (about 1.0). No R² or correlation is reported for the clinical variables.
- A population model. Nothing is person-specific. Individual differences enter only as regression targets.
- fMRI itself: a slow haemodynamic signal (about 1 Hz here) over 424 parcels.
What it means for Amadeus (inference)
- A "base" for human brain dynamics exists, and it is weak at what a model of one brain most needs. Self-supervised pretraining on thousands of hours helps (forecast R² 0.09 against 0.01 without pretraining), but the next 20 seconds of whole-brain parcel activity are mostly unpredicted. Either fMRI at this resolution is noise-limited, or the dynamics depend on things the signal does not carry: the content of thought, external input. Both point the same way. Coarse fMRI dynamics alone will not carry a person's thinking; anchors that tie activity to content are needed (Goldstein 2022, MindEye2).
- Scale comparison for the month idea. BrainLM's whole corpus is 6,700 hours across tens of thousands of people. A month of continuous recording of one person is about 720 hours, roughly a tenth of that corpus, all from one brain. This is not possible with fMRI: a person cannot stay in a scanner for a month (R1 covers what dense sampling has achieved). With wearable or implanted recording, one person's month would be a large dataset by current standards.
- Self-supervised then fine-tuned is the right recipe for the "model of one brain" that the architecture doc proposes: pretrain on the person's own unlabeled stream, then tie to meaning through labelled anchors. The evidence here is only that pretraining helps downstream tasks; nothing shows it captures a person.
- Caution on claims: the paper's framing ("interpretable", "zero-shot", "never-before-seen") is stronger than its evidence. Its numbers should be cited, not its adjectives.