Imaging Neuroscience 3, IMAG.a.20, doi 10.1162/IMAG.a.20; open access (CC BY 4.0), PMC12319872. Read for research direction R2, 4 October 2026. Provenance: papers/amadeus/kam2025_eeg_ongoing_thought.provenance.json.
What was read
- Read in full: the PMC XML converted to text: abstract, introduction, methods, results, Table 1 (probe questions), Table 2 (classification), figure captions (figures are images), discussion, limitations, conclusion, acknowledgements, data statement and the reference list.
- Not read: the online supplement (time windows of 8 and 16 s, visual and auditory dimensions, trial counts per session, CNN architecture and hyperparameters, control-analysis tables S7 to S11). Claims about those rest on the main text only.
What they did
- Design: dense sampling of few people. 7 students (4 women, 3 men; mean age 24.6), mostly from a psychology department. Each did 7 sessions of about 2 hours: 49 data sets.
- Naturalistic task. About 80 minutes per session of whatever they chose to do on a computer: reading 32%, watching videos 19%, writing or editing 14%, browsing 11%, demanding tasks such as coding or sudoku 23%, nothing 1%.
- Experience sampling. Multi-dimensional thought probes every 90 to 150 s (35 per session), each with a tone and a pop-up. Questions: what were you doing (open), then 7-point ratings of on- versus off-task, freely moving, difficulty disengaging ("sticky"), goal-directed, internal versus external focus, about yourself, about other people. Visual and auditory modality items were exploratory (supplement).
- Rest. Two 5-minute eyes-open rests per session, each followed by one retrospective report on the same dimensions (no task items).
- Recording. 32-channel wireless semi-dry EEG (Bitbrain Versatile), 256 Hz. ICA artefact removal, 2-s epochs, periodic power 4 to 30 Hz after removing the aperiodic 1/f component (FOOOF).
- Analysis.
- Ratings dichotomised: 1 to 3 against 5 to 7; a rating of 4 discarded.
- EEG from the 12 s before each probe gets that probe's label. 8 and 16 s gave the same results except two comparisons that lost significance (supplement).
- Cluster-based permutation tests across the 49 data sets, then mixed models with participant and session as random effects, and controls for the three most correlated thought dimensions.
- A shallow 1-D convolutional network on 32 channels × 27 frequencies. Within-person: 5-fold cross-validation over randomly partitioned trials. Across-person: leave one person out, then retrain on 25% of the test person's data (transfer learning) and test on the other 75%. Minority class oversampled (SMOTE) in training only. 25 repetitions; significance from 25 label permutations and a Wilcoxon test.
Main results (verified)
- Most common thought types while doing self-chosen work: freely moving, goal-oriented and internally oriented.
- Every dimension had a distinct oscillatory pattern during the task. Examples:
- off-task thought: more posterior low alpha, widespread high alpha, centro-parietal beta;
- sticky thought: more central high theta and fronto-central low beta;
- goal-oriented thought: less alpha and less high beta;
- thoughts about oneself: more frontal low alpha, widespread high alpha and beta;
- thoughts about others: more widespread low beta. All survived the participant-level and correlated-dimension controls. Age and sex did not matter.
- Task and rest patterns agree for some dimensions only. Electrode-by-frequency maps correlated r = 0.70 (sticky), 0.76 (self), 0.70 (others), 0.41 (goal-oriented). Freely moving and internal thought could not be tested at rest: the single retrospective report did not spread across both ends of the scale.
- Classification is above chance and modest. Table 2, means over 25 repetitions:
| Dimension | Within person MCC / AUC / BA | Across person (with 25% transfer) MCC / AUC / BA |
|---|---|---|
| on- vs off-task | 0.25 / 0.70 / 0.61 | 0.17 / 0.63 / 0.58 |
| internal vs external | 0.34 / 0.73 / 0.66 | 0.26 / 0.69 / 0.63 |
| freely moving | 0.22 / 0.68 / 0.60 | 0.16 / 0.63 / 0.58 |
| sticky | 0.22 / 0.69 / 0.60 | 0.14 / 0.63 / 0.57 |
| goal-oriented | 0.35 / 0.73 / 0.66 | 0.27 / 0.69 / 0.63 |
| about self | 0.34 / 0.74 / 0.67 | 0.23 / 0.68 / 0.62 |
| about others | 0.43 / 0.80 / 0.71 | 0.31 / 0.74 / 0.66 |
Within-person models beat across-person models on every dimension. The within-person standard deviations are large (MCC SD 0.12 to 0.27), so some people are decoded much better than others.
Limits
- Seven psychology students who may know the mind-wandering literature, at a desk in a lab. "Naturalistic" means self-chosen computer work, not daily life.
- The labels are the reports. Accuracy is agreement with a dichotomised 7-point rating, not with the thought itself. Middle ratings were discarded, which makes the problem easier than the real one.
- Possible leakage in the classifier tests (my reading, not raised by the authors; corrected by the main session, 4 October).
- The datapoints were "randomly partitioned" into folds, and the classifier input is "each trial".
- The paper never says whether a "trial" is one 2-s epoch or the whole 12-s window before a probe, with its epochs averaged. The window was chosen "to maximize the number of epochs needed to create a reliable average", which suggests the latter.
- If a trial is an epoch, neighbours from one probe can sit in both training and test sets, here and in the 25% transfer split, and the numbers are optimistic.
- If it is the window, the remaining risk is only correlation between nearby probes in one session.
- The paper reports no split by probe or session.
- The "across-person" test is not zero-shot: the model is retrained on a quarter of the test person's labelled data.
- Only 25 permutations for the null. Form and focus of thought are decoded, not its content (what the person was thinking about). Probes every 2 minutes may themselves change the stream of thought, as the authors note.
- Strengths: data and code are on OSF. The aperiodic correction rules out the explanation that the effects are broadband 1/f shifts.
What it means for Amadeus (inference)
- The month-of-recording design has a precedent. Wearable-class 32-channel EEG, a self-chosen activity, a probe every 2 minutes, repeated over 7 sessions per person yields labelled thought-type data. A month at a few hours a day would give one person far more probes than this whole study had per person.
- What it yields is the type of thought, not its content. On, off, self, other, sticky, goal-directed: binary, with balanced accuracy 0.6 to 0.7. That is a coarse time series of how a person's mind moves. It may characterise their dynamics, for example how often they get stuck and on what kind of activity. It does not say what they were thinking, so it cannot fill memories or beliefs.
- Person-specific calibration pays. Within-person models beat transfer on every dimension, and people differ widely in how decodable they are. A month with one person is the right shape for a person-specific decoder. The decoder still learns to predict the person's own reports, so it inherits whatever is wrong with them.
- Lessons for our evaluation: split by probe and by session, never by epoch. Keep middle ratings in the test set. Report a probe-only baseline: how well do the person's own earlier reports predict the next report?