Kurisutina

MindEye2: shared-subject models enable fMRI-to-image with 1 hour of data

ICML 2024 (PMLR 235), arXiv:2403.11207v2 (15 June 2024), 22 pages including the appendix. Read by the main session for research direction R3, 4 October 2026. Provenance: papers/amadeus/scotti2024_mindeye2.provenance.json.

What was read

  • Read in full: every page of the arXiv PDF (pdftotext -layout, two-column layout interleaved in the text): main text, tables 1–5, Limitations, Impact Statement, references, and Appendix A.1–A.15 (tables 6–11, the human preference experiments).
  • Not read: the figures as images. Reconstructions, UMAPs and the scaling curves in Fig. 5–8 were only seen through their captions. The code was not checked.

What they did

  • Data. The Natural Scenes Dataset (NSD): 8 people, each scanned for 30–40 hours (30–40 sessions) at 7 T. Each session showed 750 natural images (COCO) for 3 s each; every image was seen 3 times.
    • Images are unique to each person, except 1,000 shared images, which form the test set.
    • Input: single-trial GLMsingle betas over 13,000–18,000 visual-cortex voxels in each person's native space.
  • Model.
    • A subject-specific linear ridge layer maps each person's voxels into a shared 4096-dimensional latent space.
    • Everything after it is shared across people: an MLP backbone, then a diffusion prior into OpenCLIP ViT-bigG/14 image space, a retrieval submodule and a low-level submodule.
    • A Stable Diffusion XL "unCLIP" turns predicted CLIP embeddings into pixels. A GIT captioner predicts a caption that guides an SDXL refinement step.
  • Shared-subject training. Pretrain on 7 people; fine-tune on the 8th with limited data, taken chronologically: "1 hour" means that person's first session. Test images come from held-out sessions.

Main results (verified, from Table 1 and the appendix)

  • Full data (40 h) is state of the art on most metrics.
    • Image retrieval 98.8% and brain retrieval 98.3% among 300 candidates (chance 0.3%).
    • Two-way identification: CLIP 93.0%, Inception 95.4%.
    • Human raters picked the correct reconstruction over a random one 97.82% of the time.
  • One hour of a new person's data:
    • Image retrieval 79.0%, brain retrieval 57.4%; CLIP two-way 79.2%; Inception 81.2%.
    • Per subject, retrieval varies widely: image 94.0, 90.5, 66.9 and 64.4%; brain 77.6, 67.2, 47.0 and 37.8%.
    • Against Brain Diffuser at 1 hour, raters preferred MindEye2 only 53.01% of the time (p = 0.044).
  • Pretraining matters most at low data. Single-subject models trained from scratch on 10–30 minutes were unstable ("mode collapse"). The gap closes as data grows (Fig. 5).
  • Which people are used for pretraining matters little. Pretraining on subjects 2, 5 and 7, or on subject 5 alone, gave about the same 1-hour results as pretraining on all seven (Table 10).
  • Without pretraining, MindEye2 still beats MindEye1 (Table 6). Part of the gain is architecture: a linear first layer and the bigG target.
  • Brain-correlation check. Reconstructions are fed through an image-to-brain encoding model and compared with the measured activity. By region: r = 0.327–0.373 for refined reconstructions, 0.339–0.385 unrefined, and 0.300–0.351 with 1 hour of data (Table 3). The unrefined reconstructions score higher than the refined ones. Refinement buys naturalism at some cost in faithfulness.

Limits (the authors' and mine)

  • Authors:
    • "decoding is easily resisted by slightly moving one's head or thinking about unrelated information";
    • only natural scenes (COCO) were tested;
    • extending to mental imagery is future work, not shown.
  • Normalisation leak (the authors' own note, A.2). Voxel z-scoring used statistics from the full training set even when fewer sessions were used for training. This "may inadvertently give a small normalization advantage" to the low-data models, the 1-hour headline included.
  • "Confidently wrong" outputs (A.15.3). The strong generative prior produces natural-looking but wrong images. Raters sometimes preferred a rival's distorted output that kept true features.
  • Nearest-neighbour retrieval from a pool not containing the seen image worked poorly with bigG embeddings (A.8).
  • Perception only: seen images, 3 s each, heavy compliance (no head motion), and the scanner itself.

What it means for Amadeus (inference)

  • The individual lives in a linear map. Here everything person-specific is one ridge layer into a shared space; the rest is shared. For decoding perception that is enough, because what is decoded (the seen image) is shared by design. It is the same pattern as Wang 2025. Neural-decoding work so far treats the individual as a coordinate transform onto a shared model. A person slot that holds what differs between people needs more than that: what the person does with the shared content (Charest 2014's personally meaningful items).
  • Data scale. One hour gives usable decoding with a pretrained base; 40 hours is far better. NSD's 30–40 hours per person is the densest public perception set in this batch. The month idea (about 720 h) would be 18–24 times more per person, but no scanner allows it (R1 covers feasible modalities).
  • The confabulation warning transfers directly. A decoder with a strong generative prior fills in plausible content and is "confidently wrong". Amadeus faces the same risk: a base filling gaps in the slot. The paper's brain-correlation check (does the output, run back through an encoding model, reproduce the measured activity?) is a round-trip faithfulness test. It is the neural analogue of the formation study's link check, and worth adopting for any decoded memory: content must be traceable to the measurement, not to the prior.
  • Resistance cuts both ways for "people lie". The subject can defeat decoding by thinking of something else. Neural recording is therefore not a lie detector for a non-cooperating person. It still adds a second channel for a cooperating one, whose self-deception (as opposed to deliberate deception) is not under voluntary control. R2 covers self-deception.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.