Kurisutina

Ongoing thoughts at rest reflect functional brain organization and behavior

bioRxiv 10.1101/2025.08.18.670664, version 2 posted 13 May 2026 (CC BY-NC-ND 4.0). Published as Nature Human Behaviour (2026), doi 10.1038/s41562-026-02563-9, behind a paywall ("access: No" on the publisher page), so the published version was not read. Read for research direction R2, 4 October 2026. Provenance: papers/amadeus/ke2026_ongoing_thought_rest.provenance.json.

What was read

  • Read in full: the 38-page preprint PDF, converted with pdftotext -layout: abstract, introduction, results, figure captions 1 to 5 (figures are images), discussion and limitations, methods, data and code statements, acknowledgements, contributions and the reference list.
  • Not read: the supplement, which holds the rating distributions, the per-dimension prediction statistics (Suppl. Table 5), balanced topic accuracy per person, robustness controls and the CCA variable lists. It is not in the PDF. Where a number below says "main text", the exact value is in the supplement.
  • Version caveat: this is the last preprint before journal publication. The peer-reviewed text may differ.

What they did

  • Sample. 60 adults (34 women, mean age 22.9), two 3-hour fMRI sessions a mean of 10.9 days apart. 4 people did not return for session 2.
  • Annotated rest. 4 runs of 10 minutes, each with 8 trials. Each trial: 30 s of eyes-open rest, then 10 s of speaking aloud ("Briefly describe your thoughts just before this"), then 9 slider ratings. The ratings covered awake, external environment, future, past, self versus others, valence, images, words, and detailed and specific. That gives up to 32 sampled thoughts per person.
  • Three thought measures.
    • Ratings, z-scored within person.
    • Content: transcripts encoded with the Universal Sentence Encoder (512 dimensions).
    • Topics: 5 annotators assigned one of 9 topics: environment, body and introspection, the study's movies, social, future planning, past memory, obligations, zoning out, other. At least 3 of 5 agreed on 94.5% of thoughts.
  • Brain measure. Functional connectivity from each 30-s window (268-region atlas, 5-s lag shift, strict motion censoring).
  • Models.
    • Connectome-based predictive models: support vector regression for ratings and classification for topics, leave-one-subject-out.
    • Tests against measures that are not self-report: pupil size, a published sustained-attention network, and the sentiment of the speech.
    • The imagery model applied to an aphantasic person and their identical twin.
    • All thought models applied to 908 Human Connectome Project resting scans, then canonical correlation with 158 behavioural and trait measures.

Main results (verified)

Thought itself (self-report)

  • Stable and person-specific. A person's 9-dimension profile correlated across the two sessions at mean r = 0.62 (95% CI 0.53 to 0.70). Within-person similarity exceeded between-person similarity, for ratings (t(53) = 9.96) and for content embeddings (t(49) = 11.24).
  • More idiosyncratic as rest goes on. People became less similar to each other over the 32 trials. They became less awake and less focused on the environment, and turned more to the past and to images.
  • Main topics: the current environment (the scanner), body and feelings, and future planning. People used 2 to 9 of the topics (mean 6.4).

Brain against thought

  • Similar thoughts go with similar connectivity.
    • Between people, content similarity tracked whole-brain connectivity similarity: r = 0.096, permutation z = 4.06. Rating similarity did not (r = 0.026).
    • Within a person, trial by trial: content r = 0.069, ratings r = 0.071, both p < 0.001.
  • Decoding works for 5 of 9 rating dimensions and is "relatively modest" (the authors' words).
    • Dimensions predicted in held-out people: awake, external, future, valence, images.
    • Not predicted: past, self versus others, words, detail.
  • Topics are decoded weakly. 8-topic classification: 28.5% accuracy against chance of 12.5%, though the null is raised by topic imbalance. Balanced accuracy was 20.3%, and 36 of 38 people were above chance.
  • Checks against measures that are not self-report are weak but present.
    • The wakefulness model predicted pupil size, r = 0.128.
    • It also predicted sustained-attention network strength, r = 0.049.
    • The valence model predicted negative speech sentiment, r = −0.065, but not positive sentiment.
    • The aphantasic twin scored −0.31 on predicted imagery against +0.24 for the co-twin (p = 0.016). That is one pair.
  • Traits in an independent sample.
    • In the Human Connectome Project, model-predicted thoughts formed one canonical mode with the behaviour and trait measures: r = 0.317.
    • The mode runs from favourable (life satisfaction, emotional support) to adverse (substance use, rule-breaking, loneliness). It mirrors the known positive–negative mode of resting connectivity, with behavioural weights correlating 0.757 with the 2015 result.
    • It beat "null thought models" trained on shuffled reports: paired t = 3.39, p = 0.002.

Limits

  • Probing every 30 s shapes the stream. The authors list it. The "thoughts" are also largely about the scanner and the body.
  • What is decoded is the verbal report. It came in 10 seconds of speech, aloud in a scanner, with experimenters listening. It is the thought as reported, not the thought.
  • Accuracy is low. Content decoding is 8 coarse topics at balanced accuracy 20% against 12.5%. The connectivity–thought correlations are below 0.1. Content, in the sense of what exactly the person thought, is not recovered.
  • All models are cross-person (leave one subject out). No person-specific decoder was built, so the study says nothing about how much a dense one-person model would gain.
  • Item-level checks are missing on the trait side. The trait link uses a dataset with no thought reports, so it cannot be checked item by item. The authors call the effect small and the analysis correlational.
  • Preprint only; the raw data and transcripts are not public (transcripts withheld for privacy).

What it means for Amadeus (inference)

  • Spontaneous thought is a stable, individual variable. Within two sessions, the shape and the semantic content of a person's resting thought were more similar to themselves than to anyone else (r about 0.6 across sessions). The month-of-recording idea rests on that premise, and here it holds at the level of reports.
  • Labels are cheap and rich. Brief spoken reports at probes, embedded with a language model and coded for topic, are a workable label scheme for spontaneous thought. They give content as well as the dimension ratings that most experience sampling uses. For a month of recording this is the obvious anchor.
  • The brain adds little resolution so far. Connectivity tracks thought type and topic weakly. In this design, neural data gives "what kind of thinking, roughly" and links it to traits. It does not give the person's thoughts. The traits it predicts can be measured more cheaply by asking.
  • The thinking is shaped by the setting. A scanner produces thoughts about the scanner. A month of daily-life recording would need portable methods (see Kam 2025), and the content would differ from in-scanner content.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.