Kurisutina

The blinded Dream Catcher experiment

Citation: Wong W, Noreika V, Móró L, Revonsuo A, Windt J, Valli K, Tsuchiya N. The Dream Catcher experiment: blinded analyses failed to detect markers of dreaming consciousness in EEG spectral power. Neuroscience of Consciousness 2020(1), niaa006. Published 15 July 2020. DOI: 10.1093/nc/niaa006. Official PMC article; PMCID PMC7362719; PMID 32695475.

Reading completed 29 September 2026: complete published article supplied to PMC, all 2,193 extracted lines including Methods, results, discussion, nine figures, three tables, two numbered equations and an unnumbered model formula, references and declarations. All eight publisher supplementary documents were read in full: 2, 6, 1, 6, 1, 3, 5 and 8 pages, totaling 32 pages. Every main figure and table, all seven supplementary figures, all five supplementary tables and the supplementary silhouette equation were visually inspected. Main tables and equations were rendered from the official HTML/MathML; supplementary visuals were inspected in PDF page renders. Full article, figure assets and the complete supplement ZIP were acquired before substantive reading. The main typeset PDF was unavailable through attempted public routes; the complete official HTML/XML scientific content was available. No scientific supplement remains unread. No raw-data or code audit, experimental replication or reanalysis was performed. See papers/consciousness_connectomics/wong2020_dreamcatcher.provenance.json.

Finding and scope

An analysis team blinded to dream reports did not classify dreaming versus dreamlessness significantly above chance using selected EEG spectral features and a changing clustering procedure. The dataset comprised 54 one-minute pre-awakening recordings from nine people: three reports of dreaming and three confident reports of no experience per person. The same recordings were tested over five progressively less blinded stages.

This is an independent test of how well some proposed markers generalize under a stringent analysis protocol. It is not an exact replication of Siclari et al. (2017), and a null result in this small, selected sample does not establish that spectral markers cannot carry information about dreaming. The paper's broader philosophical claim that successful blind prediction would identify the constituents of consciousness also exceeds the experiment: prediction alone would not demonstrate causal necessity, sufficiency or identity with experience.

Participants, recordings and experience labels

Fifteen healthy Finnish-speaking volunteers were recruited. Five were excluded after an adaptation night because of sleep difficulties, unclear reports or excessive EEG artifacts; ten completed four experimental nights, and one without any recalled dreams was excluded from this analysis. The final nine participants included four men, ages 21–34, mean 27 years. Thus the sample excludes both poor laboratory sleepers/reporters and the participant who never recalled dreams.

The early-night serial-awakening procedure sampled the first 3–4 hours of sleep, usually after at least three minutes in NREM Stage 2 or 3. Across the nine participants and four nights there were 294 awakenings, averaging 8.17 per night. About 13% of awakenings involved sleep-onset REM intrusions and were excluded from this paper's NREM analysis. Recordings used 25 scalp EEG electrodes, two EOG channels and two EMG channels, sampled at 2,000 Hz. Supplement 1 supplies montage, reference and filter details. The repeated awakenings themselves alter ordinary sleep and are not representative of uninterrupted whole-night dreaming.

Participants had been instructed to give a free oral report immediately after the awakening signal, before a fresh prompt. The subsequent interview included questions about confidence, recency, content and duration. If no experience was reported, three questions probed confidence in its absence, a feeling of having experienced something without recalling content, and the last thing remembered. The detailed interview and examples are reproduced in Supplements 2–3.

Two raters categorized all reports as dreamless, white dream (experience without remembered content), uncertain or dreamful. Agreement was 94%, κ=.92. A second categorization rated dream complexity, with 83.8% agreement and weighted κ=.89. The selected cases required initial agreement on the basic report category; disagreements in the larger corpus were discussed or arbitrated. Dreamlessness meant a confident report of no experience immediately before awakening. White dreams and uncertain reports were excluded, rather than separately tested as intermediate outcomes.

Only static dreams, scores 1–4 on the perceptual-complexity scale, were selected. These lack reported temporal progression but can include coherent scenes, people, sounds and feelings. They are not synonymous with especially vague dreams: in the selected dreamful reports, 33% were rated almost as perceptually clear as waking, 19% equally clear and 4% clearer; 45% were quite or very vague. Eighty-five percent expressed confidence that the experience occurred just before awakening. Reported durations varied substantially: 22% a few seconds, 33% about half a minute, 26% about a minute and the remainder longer. A one-minute EEG segment therefore need not correspond to a homogeneous minute of the reported state.

The final 54 cases were selected for relatively clean EEG, static dream contents and matched sleep stages. Each one-minute segment contained three 20-second epochs of Stage 2 and/or Stage 3, under the older Rechtschaffen–Kales criteria. Supplementary Table S5 shows 33 Stage-2 epochs and 48 Stage-3 epochs in each report condition. Per-person Stage-2 counts differed by at most one epoch. Initial sleep-stage agreement was 76%, followed by consensus for the remainder.

Elapsed time and time in sleep before awakening were also approximately matched, but matching coarse stage labels and nonsignificant timing differences is not a demonstration that every relevant aspect of sleep depth, arousal or slow-wave physiology was equal. A timing difference with Cohen's d=.56 remained nonsignificant in this small sample. The paper's proposal that prior delta-power findings might reflect sleep depth is a plausible alternative, not a confound experimentally established or removed completely here.

What blinding meant

The Data Team designed collection and selected and coded recordings before contact with the Analysis Team. The latter initially knew the protocol, balanced class counts and available literature, but not participants' identities or report labels. A script randomized recording, pairing, participant and condition codes. After each submitted answer, the Data Team disclosed aggregate accuracy and a further level of grouping information. Per-case correctness was not disclosed at intermediate steps.

This is analytically blinded classification with later report-based scoring. It is not experience measured without reports: labels still derive from remembered experience after awakening. Nor is it five independent replications or a single frozen classifier tested on five new datasets. All steps reuse the same 54 cases, with increasingly informative constraints, aggregate feedback and altered preprocessing/features. The authors report these changes transparently. A separate preregistered analysis of remaining recordings is mentioned as future work, not as a completed result in this paper.

The initial briefing, reproduced in Supplement 4, sets out five stages and their permissible information. By the final stage, the two true condition groups are fully supplied; analysts only have to choose which group is dreaming. Even a correct answer there could not, by itself, be statistically above chance: it is a single two-alternative decision. The briefing explicitly recognizes this.

Methods and outcomes across the five steps

The analysis used modified evidence-accumulation clustering, called combination clustering. Many low-dimensional subclusterings were combined into a similarity matrix, sometimes weighted by clustering quality, and then divided into two approximately balanced groups. Literature-derived spectral directions supplied the assignment of the resulting groups to dreaming or dreamlessness. Group discovery and interpretation were therefore distinct stages: a stable partition need not correspond to the desired experience label.

Step Information available and principal procedure Reported outcome
1 Fifty-four unlabeled cases. Eight feature families screened for stability of clustering across successive 15-second segments. Selected 2,475 spectral features: 25 electrodes × 99 half-Hz bins. 29/54 decisions correct, 54%; two-sided binomial p=.68.
2 Twenty-seven same-person pairs, each known to contain one case of each report class. Reduced to 50 features, including scalp-average power, local spectral features, EMG and EOG. Clustered orientations of paired difference vectors. 16/27 pairs correct, 59%; p=.44.
3 Participant identity added. ICA components removed according to their incompatibility with known pair constraints; four forms of within-person/pooled clustering combined. 16/27 pairs correct, 59%; p=.44.
4 Two unlabeled within-person condition groups, three cases per group. Stronger component selection and condition-averaged difference vectors. 4/9 participant-level decisions correct, 44%; p=1.
5 The two whole-dataset condition groups supplied. No clustering. Three literature-inspired power measures determine which group is dreaming. One global decision, incorrect. This is not 54 independent incorrect predictions.

The reported binomial units change from cases to pairs to participants to one global decision. Repeated cases belong to only nine people, and choices are coupled by shared clustering, so the early counts should not be treated as 54 independent people or used to infer population precision without the design's dependence structure. None of the reported tests establishes equivalence to chance.

Step 1 considered conventional spectral bands, fine spectral bins, autocorrelation, permutation entropy, approximate entropy, ocular/muscular RMS and preprint-inspired regional features. The selected fine-spectrum representation had temporal consistency exceeding .9. This consistency metric is label-invariant agreement between partitions, 2 × |agreement − .5|, not accuracy for dreaming. Clustering ultimately grouped participant identity strongly: 24 of 27 true opposite-condition pairs fell in the same cluster. Stable subject differences were therefore an observed nuisance signal, not a hypothetical explanation.

Step 1 used 82,475 subclusterings, covering single features and random combinations of two through nine features. Step 2 reduced spectral dimensions and added other recording modalities, for 19 scalp-average spectral bins + 11 regional spectral features + 10 EMG + 10 EOG features. Pairs were centered and normalized, with clustering based on their common orientation rather than unknown signed condition differences. These steps are constrained unsupervised procedures informed by known class structure, not a conventional supervised decoder trained and tested on independent labeled subjects.

Steps 3–4 performed more than routine artifact cleaning. Components were selected using their agreement with the newly supplied pairing/condition structure. Step 3 retained an average of 14 of 26 components (range 5–23); Step 4 retained only four on average (range 1–6). This can remove nuisance variation, but it imposes a hypothesis about how class-relevant signals should cluster. Failure after such a transformation is not proof that the original recordings lack any useful signal.

Only Step 5 used the published 2017 Siclari features. Earlier steps used the 2014 preprint, which specified different frequency bands and regions. The final approximations were posterior 1–4 Hz power, 20–50 Hz power across the regions corresponding to Siclari's high-frequency finding, and a left-sided 0.5–4.75 Hz measure inspired by Scarpelli. Their differences were small, all absolute Cohen's d<.4, and jointly suggested the wrong orientation of the two true report groups.

Supplementary evidence and reproducibility details

Supplements 1–5 establish data collection, the translated interview, concrete dream examples, the original blinding instructions and exact sleep-stage counts. These make the reporting and selection process inspectable. They do not supply an independent physiological label of awareness.

Supplement 6 tests the clustering algorithm on a separate ECoG example from one epilepsy patient viewing visual stimuli. Four datasets use face/non-face or face/Mondrian contrasts and one or two recording locations. The clustering consistency scores were 11%, 4%, 100% and 41%. In the fourth dataset, choosing the second instead of the first principal component post hoc produced 100% consistency. These are not raw classification-accuracy percentages: the same transformed consistency metric is used. The positive result demonstrates that a favorable partition can be recovered in restricted data, while the poorer results show sensitivity to heterogeneity and chosen representation. It is not a successful post-hoc dreaming classifier, and it does not establish adequate power for this dream experiment.

Supplements 7–8 define all feature families, their channels, time windows, filters, interpolation and clustering procedures. Preprocessing and spectral estimation differ by feature family; they should not be compressed into a claim that a single standard pipeline was held fixed. For instance, Step 5's Siclari-inspired measures use a scalp current-source-density estimate from spherical-spline Laplacians, interpolating missing electrode locations. This is not cortical dipole source reconstruction. The other literature-inspired feature uses a different reference/filter/PSD procedure.

Two reporting details matter for exact reconstruction. Supplement 8's summary table lists the four Step-3 subclustering schemes again for Step 4, whereas its detailed Step-4 prose and main text describe pairwise clustering of nine condition-difference vectors. Supplement 7 refers to 28 channels in its comparison with Siclari, although the explicit EEG montage and feature counts use 25 scalp channels. These inconsistencies were recorded, not resolved by executing the code. The separate full algorithm-method paper and the OSF dataset were not opened in this reading.

Interpretation relative to Siclari and other report-reduction designs

The comparison with Siclari et al. (2017) is informative precisely because it is not identical. Wong uses nine selected participants, static early-night NREM dreams, 25 scalp EEG channels and approximate regional measurements. Siclari's principal localization experiment uses high-density EEG and source reconstruction; its prospective experiment uses individual baseline calibration and triggers awakening at extreme posterior spectral configurations. Wong did not reproduce that online trigger protocol, individualized threshold calibration or sample selection. The present failure therefore constrains unqualified claims that published spectral directions suffice for straightforward blind identification across designs; it does not erase Siclari's within-stage associations or prospective selected-epoch result.

Wong's stringent matching may reduce confounds, but it also narrows the range of dream contents and physiological differences. The authors' own simulation suggests approximately 80% power only for a large effect, d≥1.3, under a particular one-sided paired-Gaussian test with nine participants. That calculation does not establish the sensitivity of their complete clustering pipeline and is not evidence excluding smaller effects. A nonsignificant result from nine people cannot demonstrate that dreaming is absent from EEG spectra in general.

Like Frässle and Sergent, this paper is useful for separating an observed neural pattern from the operational procedure that links it to experience. Here the measured interval precedes the report and lacks an ongoing perceptual discrimination task, but remembering and describing the prior experience remains essential to the label. A confident report of no dream could still reflect loss of memory. Excluding white dreams makes the labels cleaner for this binary task but removes the separate experience-without-content-recall comparison available in Siclari.

No connectome, neural perturbation, direct connectivity metric or formal comparison between theories of consciousness is tested. The authors propose multivariate integration/connectivity measures as future alternatives; the present failure of selected predominantly univariate features is not positive evidence that integration is the correct explanation. The useful research constraint is methodological: a marker should be evaluated under preserved blinding in new data, with its target label, selection procedure, calibration requirements and unit of inference specified. Successful prediction would strengthen a marker; causal and structural sufficiency would still require additional evidence.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.