What was read
- PMC author manuscript (PMC11304553, NIHMS 2005151): abstract, introduction, results, discussion, full Methods, Extended Data Figs 1–10 legends, data and code statements. Read in full.
- NOT in the text obtained: Figs 1–4 legends, Table 1 raw metric values, Supplementary Tables 1–7. The bioRxiv preprint (56 pp) is also on disk but was not used for the numbers below.
Question
Can continuous natural language, not a choice among a few options, be decoded from non-invasive fMRI? And what does that imply for mental privacy?
Method
- Subjects: three (S1–S3) for all main analyses; seven for the cross-subject test. Personalised head cases. 3T Siemens, TR 2 s, 2.6 mm voxels.
- Training data: 82 narrative stories from The Moth Radio Hour and Modern Love, single speakers telling autobiographical stories, 15 story sessions over 16 sessions, about 16 hours per subject. Every subject listened to the same stories; the decoder learns how this subject's brain responds to meaning.
- Encoding model: GPT-1 (fine-tuned on Reddit and 240 Moth/Modern Love stories) layer-9 embeddings of each word in a 5-word context, four haemodynamic delays, L2-regularised regression per voxel; the 10,000 best cross-validated cortical voxels are used. Noise covariance estimated by held-out-story bootstrap.
- Decoder: Bayesian. A word-rate model (from auditory cortex for perceived speech, from Broca's area and sPMv for imagined speech and movies) says when words occur; the language model proposes continuations by nucleus sampling from a 6,867-word vocabulary; the encoding model scores each continuation by the likelihood of the recorded BOLD; beam search (k = 200) keeps the best. Parameters tuned on one calibration story in one subject; pipeline frozen before other subjects were analysed.
- Metrics: WER, BLEU-1, METEOR, BERTScore (used by default; a meaning-level similarity). Null: sequences generated by the same beam search with random instead of brain-based likelihoods. Ceiling: English→Mandarin→English round trip by professional translator and machine translation.
Results (with numbers)
- Perceived speech (held-out test story, 1,839 words, single trial): decoded sequences capture meaning and sometimes exact phrases. 72–82% of time-points scored above chance on BERTScore; time-points could be identified against each other with mean percentile rank 0.85–0.91. WER 0.92–0.94, so surface wording is mostly wrong: the output is a paraphrase of the gist. Naïve readers given only S3's decoded words answered 9 of 16 comprehension questions above chance.
- Regions: the parietal-temporal-occipital association region alone recovers 80–86% of whole-brain-decodable time-points, prefrontal 46–77%, the speech network only 28–59%. Predictions from different regions resemble each other, so the word-level representation is redundant across cortex, not modular.
- Imagined speech: subjects silently "told" memorised one-minute story segments. Story identified from five candidates with 100% accuracy in every subject (cross-entropy 0.23–0.83 bits vs 2.32 chance); 41–74% of time-points significantly decoded. Lower than perception because production and perception overlap only partly; an encoding model trained on imagined speech might do better.
- Silent films (Pixar shorts, no language, no instruction to narrate): 21–45% of time-points decoded, matching the films' audio descriptions for the visually impaired. One decoder trained on heard stories transfers to a non-linguistic task because semantic representations are shared.
- Attention: with two overlaid speakers, the decoder reconstructs the attended story (42–68% or 63–75% of time-points) and almost none of the unattended one (0–3%). Output follows conscious attention.
- Privacy, training: decoders trained on other subjects' data (anatomically aligned, volumetric or surface-based, five sets of source subjects) recover 1–2% or 1–5% of time-points against 65–82% within-subject. Cooperation is required to train.
- Privacy, use: subjects resisting by silently naming animals cut decoding to 0–3% of time-points, by telling a different story to 1–26%, by counting by sevens to 4–50%, against 52–57% passive listening. Restricting the decoder to one region does not defeat resistance. Cooperation is required to apply. The resistance data are withheld from public release.
- Training-data curve: scores rise "by an equal amount each time the size of the training dataset was doubled" (main text, Fig 4b), and the fraction of decodable time-points and identification rank "plateaued after 7 scanning sessions (7.5 h) and did not substantially increase up to 15 sessions (16 h)" (Extended Data Fig 7). Simulation shows the diminishing returns are not an artefact of overlapping subsamples.
- Test-data SNR: averaging repeats of the test story improves decoding only slightly.
- Spatial resolution: after Gaussian smoothing at 6–31 mm FWHM, decoding degrades gracefully; at the estimated resolution of current fNIRS systems "around 50% of the stimulus time-points could still be decoded" (Extended Data Fig 8). Encoding-model prediction actually improves with smoothing (less noise), so encoding and decoding performance are not coupled. This is a smoothing simulation of fMRI, not an fNIRS measurement.
- Error sources: model misspecification dominates. Word-level decoding correlates with concreteness (ρ = 0.14–0.27) but not with training-set word frequency; poorly decoded time-points track encoding-model failures (r = 0.22–0.58); removing fMRI data collapses the decoder to chance while resetting context only causes a brief dip (the decoder relies continually on brain data, not on its own momentum).
Limits (the paper's own and others)
- n = 3 for all quantitative claims; ranges reported, normality assumed.
- Semantic features only: "loss of specificity occurs when different word sequences with similar meanings share semantic features". Adding motor features or EEG/MEG timing is proposed.
- Everything is decoded relative to a language-model prior trained on narrative stories; the null model shows the prior alone does not produce the results, but the prior shapes the wording.
- Cross-subject transfer was tested only with anatomical alignment. Functional alignment methods were not tried, so the 1–5% figure is a floor for cross-subject decoding under this pipeline, not a ceiling on what shared decoders can do (compare [13] Anderson 2025).
- Imagined speech used memorised story segments, not the subject's own autobiographical memories.
- Authors warn that future methods may bypass the cooperation requirements, and that inaccurate outputs could still be misused.
What the brief uses it for, and whether it holds
Brief section 4 (evidence table), section 3 (neural channel: "approximate semantic reconstruction, not transcripts; cooperation required"), section 10.2 ("Decoders currently require cooperation to train and to apply [4]; that is a feature to preserve"), section 11 ("Tang's calibration plateaus at 7.5 h, not 16 h").
- Every number in the brief's row is in the paper: perceived, imagined and silent-video decoding; cross-subject 1–5% vs within-subject 65–82%; plateau at 7.5 h with no substantial gain to 16 h; about half of time-points at simulated portable resolution, and the brief correctly calls it a simulation; cooperation to train and apply; three subjects; hours of scanner time.
- One nuance the brief compresses: the main text describes a log-linear curve (equal gain per doubling), the Extended Data legend describes a plateau. "Plateau at 7.5 h" is the authors' own summary, so the brief is entitled to it, but a planner should expect small further gains, not zero.
- Calibration burden for step 3: the person-specific arm needs on the order of 7.5 h of scanning of narrative listening before the decoder is useful, on top of the elicitation sessions themselves. That matches the brief's insistence on measuring calibration burden and on comparing it with the same hours spent interviewing.
- Relevance to the person slot: the decoder recovers the semantic content of what a person is attending to or imagining, at gist level, in the wording of a language model. It does not recover episodes, provenance or salience, and it was not tested on the subject's own memories. The brief's placement of the neural channel as "semantic content of recalled material; possibly implicit weighting" (section 3) is an extrapolation the paper neither supports nor contradicts.
- The concreteness result is a small lead for the brief's provenance question: decoding is better for concrete, sensory words, and sensory detail is also what separates true from false recollection behaviourally ([11] Slotnick & Schacter).
Cross-references
- [5] Horikawa 2025 extends the same idea to recalled videos with a different generation method and also tests exclusion of the language network.
- [13] Anderson 2025 is the cross-person, zero-shot counterpart that Tang's anatomical-alignment result would seem to rule out; the two use different targets (experiential feature ratings vs word sequences) and different alignment.
- [3] Marek 2022: Tang is the within-person, many-hours design that Marek says gives precise individual maps.