PNAS 121(14):e2401959121, doi 10.1073/pnas.2401959121; open access (CC BY-NC-ND 4.0), PMC10998624. Edited by Daniel Schacter. Read for research direction R2, 4 October 2026. Provenance: papers/amadeus/kim2024_spontaneous_thought_decoding.provenance.json.
What was read
- Read in full: the PMC XML converted to text: significance statement, abstract, introduction, results, figure captions 1 to 5 (figures are images), discussion and limitations, methods (participants, procedure, acquisition), statements and the reference list.
- Not read: the SI Appendix, which holds the interview and story-making procedure, task details, preprocessing, the modelling pipeline, the quantisation, Tables S1 to S3 and Figures S1 to S13. The main text defers most methods to it, so method details below are as far as the main text gives them.
What they did
- Sample. 49 Korean adults included (57 scanned; 8 excluded for not attending, sleeping, misunderstanding or poor images); mean age 22.8, 21 women.
- Day 1, interview. A one-on-one interview produced 4 personal stories per participant from their own life, on positive and negative topics (safety, pleasure, danger, pain). Participants also read 6 "common" stories made the same way with pilot participants.
- Day 2, about a week later, fMRI.
- 5 story-reading runs of about 14 minutes each, with valence rated 3 times per story.
- 2 free-thinking runs of about 6 minutes each: think freely, and every 50.7 ± 5.6 s say a few words for what is on your mind. That makes 12 samples per person.
- After the scan, people rated their own thought words on 5 content dimensions. They re-read the stories with continuous self-relevance and valence ratings. In-scanner and post-scan valence agreed, within person r = 0.84.
- Models. Whole-brain activation patterns, quantised into 5 × 5 bins of self-relevance by valence (25 images per person), then principal component regression. Leave-one-subject-out and random-split cross-validation. Network and region importance came from virtual isolation and lesion analyses.
- Tests on independent data. The 12 free-thinking samples, plus two other labs' resting datasets (n = 90 and n = 60) with one post-rest rating of how self-relevant and how positive the thoughts were.
Main results (verified)
- Story reading: decodable, modestly.
- Self-relevance: leave-one-subject-out r = 0.32.
- Valence: r = 0.21.
- Valence decoding was better in people who reported more vivid reading (across people r = 0.37).
- Own stories against others' stories. The self-relevance model's response was higher for personal than for common stories (t48 = 10.18). Forced-choice classification of personal against common: 93.8%.
- Where. Default mode, ventral attention and frontoparietal networks carry both models.
- Self-relevance additionally draws on anterior insula, midcingulate and visual cortex.
- Valence draws on limbic regions, left temporoparietal junction and dorsomedial prefrontal cortex.
- The two weight maps overlap weakly (whole brain r = 0.12; medial prefrontal r = 0.27).
- Free thinking: barely. At the reporting time window the within-person correlation between model output and the person's later rating was r = 0.052 (self-relevance) and r = 0.050 (valence), one-tailed p of about 0.01. Nine published a-priori brain maps predicted nothing.
- Rest, other labs' data.
- n = 90: the best window (last 31 volumes, 14.3 s) gave r = 0.19 for self-relevance and r = 0.30 for valence, across people.
- n = 60, with that window fixed in advance: valence r = 0.32. Self-relevance r = −0.09, which became r = 0.22 only for an exploratory 10-volume window.
Limits
- The authors call the transfer to free thinking and rest weak. They say it does not survive correction for multiple comparisons and carries a risk of false positives.
- A confound sits under the 93.8%. Personal stories were also more familiar (0.87 against 0.74) and drew more concentration (0.78 against 0.69). The authors name this. The forced-choice result shows that own-life text can be told from others' text. It does not show that the self-relevance of a thought can be read.
- Window choice for rest was partly post hoc. The best window in the n = 90 set and the 10-volume window in the n = 60 set.
- Models are cross-person, and the labels are thin. Each person contributes 12 free-thinking samples, rated after the scan.
- The authors point to idiosyncrasy. Representations of spontaneous thought may be "highly complex and idiosyncratic". Their earlier work found valence representations becoming more idiosyncratic as topics become more self-relevant. They suggest extensive sampling of a few people.
What it means for Amadeus (inference)
- A template for deliberative recall under recording. Interview the person, turn their own life into stimuli, record them reading or recalling it, and collect their ratings. Deliberative recall adds the person's recall and reflection to the same loop. Holding stimulus content partly constant (personal against common stories) is what made the self-relevance model checkable.
- The fields the slot wants from neural data are hard to read today. Brief 5.1 wants salience, affect intensity and self-relevance per memory. On a person's own spontaneous thoughts, the transfer correlations are about 0.05 within person. Group models do not read these fields usefully during free thought.
- Small-N dense sampling is the open bet. The paper itself points there. Whether person-specific models do better is untested here. That is exactly what a month of recording one person would test.
- Self-report agreed with itself here (in-scanner against post-scan valence, r = 0.84). For affect labels the "say" channel is reliable in this setting. The question for neural data is what it adds beyond that.