What was read
- Full open-access article (nature.com PDF): abstract, introduction, results, discussion, methods, data and code availability, references. Read in full.
- NOT read: Supplementary Figs 1–17 and Tables 1–2 (feature rating instructions, sentence list, word clouds), referenced but not in the PDF.
Question
Do self-generated autobiographical mental images and externally driven sentence comprehension share cortical representations? Operationally: can a decoder trained on other people reading sentences reconstruct a person's own ratings of their imagined autobiographical experiences, zero-shot across people and tasks?
Method
- Imagery data (reanalysed from Anderson et al. 2020): 50 healthy people, 25 young (24 ± 3) and 25 elderly (73 ± 7). Outside the scanner, each imagined their own experience of 20 generic scenarios (resting, reading, writing, bathing, cooking, housework, exercising, internet, telephoning, driving, shopping, movie, museum, restaurant, barbecue, party, dancing, wedding, funeral, festival), gave a brief verbal description, then rated each on 20 experiential features (sensory, motor, affective, social, cognitive, spatiotemporal; 0–6), plus whether it was a real event (means 5.5–5.6 of 6) and how vivid (4.6–5.2). In the scanner (3T Prisma, 2 mm voxels, TR 2.5 s, 639 volumes, about 27 minutes), generic written prompts ("a wedding scenario") cued 7.5 s of re-imagining, five randomised repetitions; four post-onset volumes averaged, then averaged across repetitions to one volume per scenario.
- Sentence data: 14 different participants (mean age 32.5) reading 240 simple sentences ("the child broke the glass in the restaurant") at a 3T GE scanner; sentence meaning modelled as the sum of crowd-sourced ratings of its content words on the same 20 features.
- Common space: all data averaged into the Schaefer-1000 parcellation, deliberately blurring individual anatomy; default mode network subsystems defined as Core (82 parcels), fronto-temporal FT (66), medial temporal MT (27).
- Decoder: ridge regression (λ = 1e5, chosen by leave-one-participant-out cross-validation within the sentence data) from 1000 parcels to each feature, fitted on all 14 sentence participants, then applied unchanged to the imagery data. Evaluation: Spearman correlation between reconstructed and self-rated features across the 20 scenarios per person; two-alternative forced-choice discrimination of participant pairs (400-entry vectors) and of scenario pairs; permutation tests; FDR correction. Comparison decoder using GPT-2 layer-16 embeddings of the verbal descriptions.
Results (with numbers)
- Feature reconstruction: 11 of 20 features reconstructed above chance across participants (FDR < 0.05), in order: Social, Speech, Communication, Head, Auditory, Pleasant, Music, Lower-limb, Bright, Motion, Touch. Within DMN subsystems 15 features in total; MT-DMN adds Path, Colour, Taste, Landmark, Unpleasant.
- Scenario reconstruction: 12 of 20 scenarios above chance, mean Spearman 0.06–0.22 (Telephoning 0.22, Funeral 0.20, Cooking 0.18). Within-person scenario-pair discrimination 64%, significant in 25 of 50 participants; social scenarios most discriminable.
- Person specificity: participant pairs told apart from the reconstructed 20×20 rating matrices at 70% (chance 50%, p < 1e-4); elderly 77%, young 67%. The Social feature alone gives 76%. Per subsystem: Core 76%, FT 69%, MT 75%. Single-scenario vectors almost never discriminate participants (only Exercising, 66%).
- Which features matter: variance partitioning shows seven features uniquely predict imagery fMRI (Social across all three subsystems, Path, Colour and Lower-limb in MT-DMN, Unpleasant and Motion in FT-DMN, Communication in Core-DMN); adding Time accounts for the rest of the sentence data. The other features are not separable in the sentence data used for training.
- GPT-2 comparison: reconstructing 1024 GPT-2 features of the verbal descriptions gives the same scenario-structure recovery (RSA), the same 64% scenario-pair discrimination and similar participant discrimination (Core 74%, FT 64%, MT 73%). Neither model has an advantage; parcel averaging does not hurt.
- The decoding mapping, and one fitted across both datasets, are released publicly (OSF D4KU8).
Authors' conclusions: imagined autobiographical experience and sentence meaning share cortical areas and "at least some representational codes" in all three DMN subsystems; this is "proof of principle that idiosyncratic features of self-generated mental states can be extracted from fMRI data without first tailoring a decoding model to those specific individuals"; "future models will also need to capture the vivid autonoetic recollections that distinguish episodic memories from conceptual processing".
Limits
- Features, not episodes: the readout is 20 experiential dimensions (how social, how much speech, colour, path). Nothing about who, where, when or what happened; even the verbal descriptions were only used to build the GPT-2 comparison.
- Person discrimination is pairwise and pooled over all 20 scenarios; single scenarios do not identify the person. 70% pairwise is weak evidence of individuality, and it is stronger in the elderly group.
- Averaged fMRI: five repetitions per scenario, four volumes each, then 1000-parcel averaging. No single-trial results.
- Generic prompts read in the scanner may contribute prompt semantics; the individual-differences analysis is the control, not a separate baseline task.
- Only seven or eight of the 20 features are actually separable in the training data; coverage is limited by the sentence set.
- Supplementary material not read here.
What the brief uses it for, and whether it holds
Brief section 4: fifty people, twenty scenarios, decoder trained on fourteen other people reading sentences, 11 of 20 features, participant pairs at 70% vs 50%, zero-shot across people and tasks; limits: features not episodes, five-repetition averaging, modest pairwise discrimination. Section 3: zero-shot decoding "shows the channel reaches properties of experience, not episodes"; "shared decoders may reduce [calibration burden]". Section 9, step 3: a shared zero-shot decoder arm at matched calibration hours. Section 11: added in the second review round, replacing "7.5 h" with a measured calibration burden.
- Every number and every limit in the brief's row is in the paper.
- Two things the paper adds to step 3 that the brief could state now. First, the calibration burden of the shared arm is already measured in one configuration: zero decoder-fitting hours for the participant, about half an hour of scanning, plus the pre-scan description and rating session. Second, the decoder is public, so the shared arm can be piloted without a partner lab fitting anything.
- Ceiling on what the shared arm can deliver, from this paper: trait-like properties of a person's experiences pooled across many episodes (this person's memories are unusually social, or unusually unpleasant), not the identity of an episode. That maps onto the brief's "association-and-salience graph" only at the level of global weights, not links. Recovery of per-episode content would need per-person decoders of the [4]/[5] kind or much richer training data, which the authors themselves say.
- For the provenance hypothesis (section 3 box): the features most reliably reconstructed (Social, Speech, Communication) are the ones with the least to say about whether a recollection is genuine; sensory features (Colour, Taste, Path) come only from the small MT-DMN subsystem. So the neural channel's contribution to provenance is not demonstrated here and is probably the harder of the two hypotheses.
- Elicitation protocol note for step 1: the scenario-plus-description-plus-ratings procedure is a cheap, scanner-free way to build the target vectors the neural arm would be scored against. It is also the kind of structured recognition-and-rating task that the brief lists under behavioural elicitation, and it took one session.
Cross-references
- [4] Tang 2023: within-person decoders at 7.5–16 h; Anderson is the cross-person alternative at zero fitting hours with a far coarser readout.
- [5] Horikawa 2025: also perception-trained decoders transferred to imagery, but within person and at the level of scene descriptions.
- [6] Bonnici 2012: within-person identification of which of three real episodes is being recalled; the complement of Anderson's cross-person, feature-level result.
- [12] Park 2026: the behavioural baseline the shared neural arm must be compared against at matched participant time.