Biological Psychiatry 94(6):445–453, doi 10.1016/j.biopsych.2023.01.020; NIH author manuscript, PMC10394110 (NIHMS1884002). Erratum 2024 (96(4):322, PMC11269009). Read for research direction R1, 4 October 2026. Provenance: papers/amadeus/xiao2023_depression_decoding.provenance.json.
Why this paper and not Sani 2018 or Kirkby 2018
The two landmark multi-day mood-decoding studies in epilepsy patients were not readable: Sani et al. 2018 (Nature Biotechnology) is paywalled with no open copy, and Kirkby et al. 2018 (Cell) is behind a Cloudflare bot check. This paper uses the same design (continuous intracranial recording, sparse self-reports as labels, person-specific decoders) and is open. It cites both as its precedents.
What was read
- Read in full: the PMC author manuscript (NCBI efetch XML converted to text): abstract, introduction, Methods and Materials, Results, Discussion, acknowledgements and disclosures, all figure captions, the key resources table and references. Also the erratum, which only adds a missing financial disclosure (royalties from nView and OCDscales for one author).
- Not read: the Supplement (feature-extraction details, the region-selection and permutation methods, Figures S1–S6, which hold the detrending, 5-fold and resting-state control results). The main text reports those results; their figures were not seen. The medRxiv preprint was not read.
Design
- Who: three people with severe treatment-resistant depression and no epilepsy, in an early-feasibility trial of individualised DBS guided by recordings (NCT03437928). Baylor IRB and FDA investigational device exemption. Each received permanent DBS leads plus temporary stereo-EEG (sEEG) electrodes in four prefrontal regions targeted for depression: anterior cingulate (ACC), dorsolateral prefrontal, orbitofrontal and ventromedial prefrontal cortex.
- Duration: nine days of inpatient monitoring after implantation.
- Labels: the Computerized Adaptive Test for Depression Inventory (CAT-DI): about 12 items chosen adaptively from 389, taking 1–2 minutes, with test–retest reliability r = 0.92. It was given repeatedly, giving 36, 27 and 47 labelled time points in the three patients over nine days (about 3–5 a day).
- Neural features: sEEG at 2 kHz, recorded during each assessment; bipolar re-referencing; Hilbert power in six bands from delta to high gamma. About 400 features per patient. No stimulation during recordings.
- Model: LASSO regression with automatic selection of a single region, using training data only. Leave-one-out cross-validation; permutation test on normalised RMSE.
Main results (verified)
- Severity fluctuated and fell over the stay: CAT-DI means and standard deviations were 78.9 (6.6), 63.0 (4.6) and 64.5 (14.1). The authors list stimulation, contact with the team "and/or other effects" as causes of the decline.
- A consistent spectral direction: among significant features, low-frequency power (delta to beta) correlated positively with severity (88.5–99.2% of them). High-frequency power (gamma, high gamma) correlated negatively (92.9–100%).
- Decoding held-out time points within each person:
- predicted against measured r = 0.66, 0.93 and 0.68; normalised RMSE 0.75 (p < 0.01), 0.37 (p < 10⁻⁴) and 0.72 (p < 10⁻³);
- pooled standardised r = 0.73;
- ACC was selected in all folds for patients 1 and 2 and in 41 of 47 folds for patient 3.
- Controls:
- Time trend: neural features predicted severity residuals after removing a linear time trend, all p < 0.05.
- Other validation: 5-fold cross-validation was also significant. Without region selection, decoding was still above chance but less accurate.
- The best features differed by person: ACC beta was the most predictive feature in patient 1, ACC high gamma in patients 2 and 3, with informative features spread across prefrontal regions "in an individual-specific manner".
- Is it the questionnaire or the mood? In patient 3 only, a 1-minute rest recording after each assessment decoded severity about as well. Its correlation pattern matched the during-assessment pattern (r = 0.96), and so did its feature predictiveness (r = 0.85).
- The authors' own conclusion about labels: mood decoding "relies on administering behavioral assessments, a process that suffers from subjectivity and places a high burden on patients to frequently and accurately report their mental state". They contrast this with motor decoders, which have objective, temporally dense labels, and suggest video, audio, facial expression and speech analysis as future label streams.
Limits
- Very small samples. Three patients, and 27–47 labels each, against about 400 features. Leave-one-out on autocorrelated time series within nine days can be optimistic. Detrending and 5-fold checks help, but there is no held-out later period (no forward-in-time test).
- Label and recording are simultaneous. Labels were recorded during the self-report itself; the rest-period check rests on one patient.
- Unusual conditions. Days 1–9 after brain surgery, in a monitoring unit, with severity drifting downwards.
- Supplementary figures unread.
What it means for Amadeus
- Sparse self-reports can label multi-day intracranial data well enough to decode an internal state (verified, n = 3). About 30–50 one-to-two-minute self-reports over nine days were enough to fit person-specific decoders of a slow state (depression severity) that tracked held-out reports at r 0.66–0.93. This is the strongest available support for brief 6.7's idea of experience sampling as the labelling layer: for slow affective states, in intracranial recordings, within a person.
- What it does not show (inference):
- Decoding of content: what the person was thinking about, remembering or imagining.
- Anything non-invasive. The signal is prefrontal sEEG.
- A forward-in-time test. A month-long design should hold out whole later days or weeks, not interleaved points.
- Person-specificity again (verified): the predictive features differed between people even within one diagnosis and one region. That supports per-person models (P3: identify this person's variables), and warns that cross-person decoders will lose the individual part.
- Report burden is the bottleneck the authors name (verified). Frequent, accurate self-report is costly and subjective. For Amadeus that means the labelling budget (how many probes per day, of what length) is a first-class design parameter, and passive streams (audio, video, speech, physiology) should carry as much of the labelling as they can.