What was read
- Full text from PubMed Central (PMC12588295): abstract, introduction, results, discussion, all Materials and Methods, statistical analysis, Fig 1–4 legends and Table 1. Read in full.
- NOT read: supplementary figures S1–S14 (referenced but not in the XML).
Question
Can structured descriptions of visual mental content (objects, actions, and their relations, not just word lists) be generated from fMRI, both for what a person is watching and for what they are recalling, and does this require the language network?
Method
- Six subjects (S1–S6), all native Japanese speakers, non-native English speakers (S6 with limited English); scanned over about 6 months; about 17.1 hours per subject in total: 11.8 h training (2,180 unique short videos, each viewed once, silent, no fixation), 1.8 h test (72 videos, five repetitions each), 3.6 h imagery (30 runs). Parameters were set on a pilot with S1. 3T Prisma, TR 1 s, 2 mm voxels, multiband 6.
- Captions: 20 crowd-sourced English captions per video (MS-COCO style), proofread with ChatGPT. Subjects also rated, in-scanner, how well sample captions matched what they had perceived.
- Stage 1, brain-to-features: ridge regression from up to 50,000 whole-brain voxels to DeBERTa-large layer-wise semantic features of the captions (42 language models compared; DeBERTa-large chosen by encoding performance). Data samples are block averages; test and imagery samples are averaged over five repetitions unless stated.
- Stage 2, features-to-text: iterative optimisation starting from a single unknown token. Mask one to three words or insert masks, let RoBERTa-large (masked language model) propose fillers, score candidates by mean layer-wise correlation with the brain-decoded features with a length penalty (α = 0.1), keep the top five, repeat 100 times, five restarts. No caption database, no trained captioning network.
- Imagery: subjects saw a verbal cue (caption) for one of the 72 test videos, pressed a button, closed their eyes and replayed the video mentally for the video's duration, then watched it and rated accuracy and vividness. They had practised the video–caption pairings beforehand. Decoders were trained on perception only and applied unchanged to imagery.
- Evaluation: discriminability (score for correct captions minus mean over 2,179 wrong videos) on feature correlation, BLEU-4, METEOR, ROUGE-L, CIDEr, BERTScore; identification among 2 to 100 candidates; word-order shuffling controls; comparison with database search (4.1 M captions) and with a nonlinear captioning model (CLIP→GIT); language-network localiser and ablation.
Results (with numbers)
- Viewed content: descriptions evolve from fragments into coherent sentences that capture objects, actions and their changes. Discriminability significant on all metrics for all subjects (P < 0.01 FDR). Identification among 100 candidates about 50% for all subjects (chance 1%) using feature correlation (Fig 2E).
- Structure is real, not imposed: shuffling all words or only nouns lowers discriminability significantly, even for the most fluent shuffles; brain-decoded features correlate best with the original order (top of all all-word shuffles, top 0.001% of noun shuffles); decoders trained on shuffled-caption features produce incoherent output; the optimiser reconstructs reference captions exactly (correlation 1.0) from model-derived features. An untrained masked language model does not produce comparable output.
- Beats alternatives: higher discriminability than caption-database search and than the nonlinear captioning baseline on both this video dataset and the Natural Scenes Dataset (images); generated descriptions align with the captions each subject rated as most consistent with their own perception; better brain-aligned language models give better text.
- Regions: semantic encoding model wins over a video-recognition model from the anterior half of category-selective visual cortex onward into parietal and frontal cortex and the language network. Decoding from voxels best predicted by the semantic model approaches whole-brain; the language network alone gives low performance; ablating the language network leaves almost 50% identification among 100 for viewing, with the shuffling effect intact and intelligible descriptions.
- Recalled content: perception-trained decoders applied to imagery produce descriptions matching the recalled video; "proficient subjects achieving nearly 40% accuracy in identifying recalled videos from 100 candidates" (chance 1%), with clear individual variation (Fig 4B). Shuffling still hurts; excluding the language network reduces accuracy slightly, not substantially. Semantic features generalise from perception to imagery better than visual (TimeSformer) or visuo-semantic (CLIP) features at every layer. Decoding from the preparation period (reading the cue) is near chance for most subjects, so the imagery itself carries the signal. Single-trial imagery decoding gives "reasonable identification accuracy" (Fig 4F).
Author's own limits and framing
- Captions were written by third-party annotators about visual content, so outputs are "predominantly concrete and rarely reflected abstract dimensions such as impressions and emotions"; training on the subjects' own reports "might yield even closer alignment".
- Natural web videos: cannot tell which relational structures are captured or whether atypical scenes ("a man bites a dog") would decode; possible bias toward typical scene structure from model priors, training data or stimulus selection.
- Verbal cues may bleed into the imagery period through the slow haemodynamic response; the preparation-period control mitigates but does not remove this. Spontaneous imagery (mind-wandering, dreams) with verbal reports as reference is the proposed next test.
- The method is "an interpretive interface rather than a literal reconstruction of mental content": outputs reflect brain-derived information and priors from the language models, annotation language and style, and stimulus properties. "How much of the decoded output truly originates in the brain, and how much reflects the constraints of our tools?"
- Privacy: intensive data collection from willing participants currently ensures consent, but "advances in interindividual alignment technology could reduce this requirement"; language-model biases could distort outputs.
Other limits relevant here
- Six subjects; parameters tuned on one of them who saw the stimuli repeatedly.
- Recall is of short, recently viewed, rehearsed video clips, cued by a caption, with five averaged repetitions in the main analysis. This is far from spontaneous autobiographical recollection of a decades-old episode.
- About 17 hours of scanning per subject, over months.
What the brief uses it for, and whether it holds
Brief section 4: viewing-trained decoders describe recalled videos not used in training; roughly 50% among 100 for viewing, nearly 40% for the better participants in recall, individual variation; performance largely survives excluding the language network. Limits given: candidate identification, six participants, controlled recently viewed material. Section 11 records the correction from "50% recall" to "50% viewing, nearly 40% recall".
- All numbers and limits hold as stated. The 72 recalled videos were excluded from decoder training. The language-network result is as described for viewing; for recall the paper says "slightly, but not substantially, reduced".
- Two things the brief's row omits that matter for its own plan. First, calibration burden: about 17 h per subject, comparable to Tang's 16 h, so the brief's "calibration burden to be measured" (section 3) already has two data points of the same order. Second, the author's "interpretive interface" caveat is the strongest statement in the evidence base of the brief's own point (5.3, 6.4) that a decoded description is a translation through priors, not a readout; the brief could cite it for that.
- For the neural channel's hypothesised role ("semantic content of recalled material; possibly implicit weighting and provenance"): this paper delivers the first (semantic content of a recalled visual scene, concrete, third-person framing) and nothing on the second or third. Emotion and impression are explicitly absent from the outputs because the annotation frame excluded them.
- For the shared-decoder arm of step 3: Horikawa's privacy paragraph points to inter-individual alignment work (refs 68, 69) as the route that would cut per-person calibration, the same direction as [13] Anderson.
Cross-references
- [4] Tang 2023: autoregressive, perception-trained, language-based; Horikawa's decoders are trained on non-linguistic video and generate by bidirectional masked optimisation; both find gist-level, paraphrased output and both need many hours per subject.
- [6] Bonnici 2012: the only paper in the base where the recalled material is the person's own autobiographical memory rather than a lab stimulus.
- [13] Anderson 2025: the zero-shot cross-person alternative.