Kurisutina

Adversarial testing of IIT and GNWT

Citation: Cogitate Consortium, Ferrante, O., Gorska-Klimowska, U., Henin, S., et al. Adversarial testing of global neuronal workspace and integrated information theories of consciousness. Nature 642, 133–142. Published 30 April 2025; accepted 11 March 2025. DOI: 10.1038/s41586-025-08888-1. Peer-reviewed, CC BY 4.0.

Reading and provenance

Read on 29 September 2026 from the complete PMC article, the published article rather than the 2023 preprint. Read all 3,599 lines of the archived article text: front matter, abstract, main text, discussion, full Methods, acknowledgments, declarations, references, data/code statements, all figure/table captions and associated-data text. Visually inspected all four main figures, all seven Extended Data figures, and all three Extended Data tables, separately archived as 14 images. Bibliographic entries were read as part of the article; their referenced papers were not opened or used as independently read evidence.

The complete 135-page Supplementary Information and eight-page Reporting Summary were also read. Reading was split across agents: the consciousness worker read physical PDF pages 1–68, all 37 figures and Tables 1–8 in that portion, plus all eight reporting pages; the structure/dynamics worker read pages 69–135, Figures 38–55, Tables 9–26 and the continued Figure 37 caption. This covers all supplementary prose, equations, table entries, captions, references, methods, deviations and separately authored theory discussions. All 55 supplementary figures and 26 supplementary tables were visually inspected. The scanned Reporting Summary was read visually in full. The scientific main article, Supplementary Information and Reporting Summary are therefore fully covered; this does not mean the raw data or analysis code were independently rerun.

Initial PMC/Springer retrieval failures were resolved through the official Springer supplementary-information CDN and reporting-summary CDN. The peer-review file and experimental video were not retrieved or read and are outside this reading claim. They are not used as evidence. The supplementary PDF's creation metadata is 24 October 2025, after original publication; its exact retrieved bytes are hashed, and metadata alone is not treated as proof of a substantive revision.

Archive: papers/consciousness_connectomics/cogitate2025_adversarial.html, .txt, _image1.jpg through _image14.jpg, _supplement.pdf, _supplement.txt and _reporting.pdf; source URLs, hashes, exact reading scope and historical access failures are recorded in cogitate2025_adversarial.provenance.json. Failed responses remain clearly named .access_failure.html and are not scientific PDFs.

Question and what was actually tested

The study compares specific agreed biological predictions of integrated information theory (IIT) and global neuronal workspace theory (GNWT) concerning the location, persistence and interregional coordination of visual conscious content. It does not directly calculate IIT's intrinsic integrated information, implement the full computational core of GNWT, manipulate consciousness on/off, or test artificial systems.

The principal distinction is between posterior cortical content and prefrontal/global broadcasting. All participants saw clear, attended stimuli. The experiment asks whether mechanisms that the theories predict should accompany such experiences are present. It does not establish that every measured representation is unique to consciousness, because there is no matched unconscious-stimulus condition. A missing predicted necessary feature is potentially informative even under that limitation; a positive decoding result alone does not identify consciousness.

The theory proponents and a theory-neutral consortium agreed predictions, regions, outcomes and interpretation before testing. Teams collected data across laboratories and three modalities. This makes the prior commitments unusually useful, but interpretation still depends on auxiliary assumptions connecting the theories to particular measurements.

Design and measurement

  • Recruited 256 people: 120 fMRI, 102 MEG, 34 intracranial EEG (iEEG). These are not the final sample sizes for every analysis. Final confirmatory fMRI used 73 healthy participants and MEG 65; separate optimization samples comprised 35 and 32. Five MEG and 12 fMRI participants were excluded. Thirty-two iEEG patients had complete data, with roughly 28–31 contributing to individual analyses. Three iEEG patients below the predefined behavioral criterion were nevertheless retained; the main Methods disclose this departure.
  • Four categories (faces, objects, letters, false fonts), 20 identities per category, three orientations, and three durations (0.5, 1.0, 1.5 seconds). Participants detected rare designated identities. Categories alternated between task relevant and irrelevant; orientation and duration remained irrelevant. A separate 39-person surprise-recognition experiment supported visibility of irrelevant stimuli. Faces and false fonts were better remembered when relevant; objects and letters were similar across relevance conditions. This is retrospective memory evidence in another sample, not trial-by-trial confirmation of conscious experience in the main recordings.
  • Task-irrelevant nontarget trials reduce report/action confounds without eliminating all task engagement. Cross-task classifiers test information that generalizes across relevance conditions.
  • iEEG used high-gamma power (70–150 Hz); MEG included source-reconstructed signals; fMRI used spatial patterns and seed connectivity. These modalities have different scales and different limits. Most iEEG multivariate decoding and representational-similarity analyses pooled electrodes across patients into a super-participant, whereas interregional connectivity required within-patient electrode pairs. A pooled decodable representation is not a fully observed individual brain's code.
  • Analyses included linear support-vector classification, cross-temporal generalization, linear mixed models of response duration, representational similarity analysis (RSA), preregistered pairwise phase consistency (PPC), and exploratory amplitude-based dynamic functional connectivity (DFC). DFC used Gaussian-copula mutual information. Statistical procedures include permutation/cluster corrections and Bayes factors; analysis changes are disclosed in Methods and Supplement section 14; these departures are summarized below.

Predictions and results

1. Where is conscious content decodable?

Commitment: GNWT required cross-task category decoding and orientation decoding in prefrontal regions. IIT expected maximal decoding posteriorly; adding PFC should not improve posterior decoding. The decoding prediction was non-critical for IIT, whose information notion concerns intrinsic causal structure, not an outside decoder's accuracy.

Finding: Category was decodable in posterior and prefrontal regions. In iEEG, face-versus-object cross-task decoding exceeded 95% posteriorly and reached approximately 70% in PFC, with the PFC representation much briefer (roughly 0.2–0.4 seconds). fMRI showed extensive posterior decoding and more restricted, weaker middle/inferior frontal decoding. MEG also decoded category in both.

Orientation was robust posteriorly. Prefrontal face orientation was absent in iEEG/fMRI, with Bayesian evidence supporting a null result in several analyses, but weakly above chance in MEG (approximately 35% versus a 33% baseline). MEG leakage controls gave mixed evidence; the posterior-to-anterior gradient is compatible with leakage. Consequently the study's integrated verdict is an inconclusive/partial challenge concerning GNWT orientation content, not a categorical demonstration that all PFC orientation information is absent.

Adding PFC did not improve iEEG or MEG category/orientation decoding and sometimes reduced it. This is compatible with IIT, but the paper explicitly says it is also compatible with GNWT if PFC broadcasts information without adding new information. A small non-critical fMRI advantage (roughly 1–2 percentage points) was present when the regions were combined. Therefore “PFC adds no decoder accuracy” does not demonstrate that PFC is causally dispensable.

2. How is content maintained through time?

Commitment: IIT predicted sustained content-specific posterior activity/representation during perception. GNWT predicted prefrontal ignition 0.3–0.5 seconds after both stimulus onset and offset, with an intervening latent state. For GNWT, activation rather than content-specific offset RSA was the primary critical measure.

Finding: 25 of 657 posterior electrodes matched the sustained duration model in the main high-gamma analysis: 12 non-category-specific and 13 category-specific. Eight of 53 face-selective electrodes matched the sustained pattern. A stricter supplementary single-trial tracking analysis identified 15 of 657 posterior electrodes (Table 10); model fit and robust sustained tracking are related but different counts. Posterior persistence exists, but is sparse among sampled electrodes; most face-selective sites responded transiently.

In PFC, 99 electrodes showed nonselective onset responses and 24 category-selective onset responses. None of 655 prefrontal electrodes in the task-irrelevant high-gamma analysis matched the preregistered onset-plus-offset pattern. The same model could identify such patterns in visual areas. An exploratory unrestricted-duration analysis identified one inferior-frontal electrode with earlier-than-predicted onset/offset responses (about 0.15 seconds). The supplementary analyses make the signal qualification essential: one task-irrelevant PFC ERP electrode fit the GNWT temporal model (Table 5, Figure 32), and task-relevant data contained three ERP electrodes and one high-gamma electrode with that model fit. Two prefrontal MEG parcels also showed a GNWT-like alpha pattern, but this was not robust to task relevance, delayed windows or alternative sustained-activity controls; leakage from stronger posterior signals remained possible. A separate supplementary control requiring onset/offset responses in both relevance conditions found nine PFC onset electrodes and zero offset electrodes (Table 9), a different selection rule from the 123 main onset candidates. The defensible conclusion is a substantial challenge to the predicted robust prefrontal pattern, not absolute absence of every onset/offset-like response in every signal.

Posterior iEEG RSA showed sustained category and object-identity information, fitting the IIT temporal model. Orientation information was not sustained, despite successful posterior orientation decoding. PFC showed transient category content but no reliable identity content or offset reinstatement. MEG did not consistently support either theory's representation-maintenance pattern. IIT passed the predefined minimum content criterion, but did not explain persistence for every phenomenological dimension tested. Supplementary direct model comparisons are stricter than a significant correlation with one template: in the base posterior RSA analysis, task-irrelevant face/object category and object identity significantly favored IIT over GNWT. Many other individually significant IIT-template correlations did not show significant direct preference. IIT-shaped PFC correlations do not satisfy the theory’s posterior-location prediction.

3. What interregional coordination accompanies perception?

Commitment: IIT predicted sustained content-specific synchronization linking early visual cortex (V1/V2) and category-selective posterior regions. GNWT predicted brief later coordination between category-selective regions and PFC.

Preregistered PPC result: Neither theory received the expected phase-synchrony pattern. Early low-frequency effects were brief and largely explained by stimulus-evoked responses. No sustained posterior gamma synchrony supported the IIT prediction, and no expected content-selective prefrontal PPC supported GNWT.

Exploratory DFC result: In the main task-irrelevant analysis, removing the evoked response left transient category-selective prefrontal amplitude connectivity, consistent with GNWT's transient prediction; posterior effects did not establish IIT's predicted sustained interaction. This must not become a universal absence claim: extended pooled/task-relevant iEEG analyses found face-selective coupling with both V1/V2 and PFC extending to approximately one second, with much of it surviving evoked-response regression (Supplement section 8). MEG effects were mainly within the first 0.5 seconds. Task condition, modality and connectivity definition therefore matter.

fMRI seeded from fusiform face area found content-selective connectivity with early visual, intraparietal and inferior frontal cortex only when relevance conditions were pooled; neither separately tested condition produced significant corrected clusters. The object-selective LOC seed was null. The exploratory status, pooling and changed connectivity metric matter: the preregistered GNWT phase-synchronization test did not simply pass.

Coverage limit: Posterior connectivity relied on only 12 V1/V2 electrodes, from four patients, versus 472 PFC electrodes in the corresponding overall coverage. The main total n = 256 should not be used to imply uniform sensitivity for each null test.

Supplementary controls, deviations and reporting

The behavioural and eye-movement sections confirm generally high task performance while documenting modality/category differences and lower, more variable iEEG performance. Fixations remained concentrated near the stimulus. Eye movements were not identical across conditions: blinks, saccades and pupil size exhibited some effects of time, relevance and category. These controls narrow possible confounds without independently proving consciousness on every trial. The surprise-memory evidence likewise supports visibility without removing the distinction between perception, memory and report.

The optimization-versus-replication comparison was broadly similar for category/orientation decoding, but not a uniform replication of every analysis. The GNWT-like prefrontal alpha fit was absent in the optimization sample; corrected fMRI gPPI clusters emerged only in the larger replication sample with relevance conditions pooled. The Reporting Summary's broad replication statement should be understood against these detailed differences. iEEG did not have a separate replication sample.

Supplement section 14 records material deviations and additions: three low-performing iEEG patients were retained; RSA changed from trial splitting to all-pairs correlations; fMRI gPPI pooled relevant/irrelevant conditions and replaced planned FDR testing with cluster permutation; amplitude DFC was added following the phase-connectivity null results; a requirement for sustained category-selective electrodes was relaxed because V1/V2 sampling was sparse; and fMRI native-space analyses were replaced by standard-space analyses. These choices can be informative but should not inherit the evidential status of every original prospective commitment. The ROI list was finalized in June 2022 after optimization data only, before analysis of the held-out two-thirds.

Exploratory putative-NCC analyses found predominantly posterior responses alongside some inferior/middle/orbital frontal responses or decodable information and widespread deactivation. They identify candidate correlates conditional on the contrasts, not a causally sufficient minimal substrate. Section 11's joint discussion and Sections 12–13's separate theorist discussions distinguish shared results from disputed interpretation. Statements by Boly, Koch and Tononi about posterior sufficiency, and Dehaene's interpretation involving attention shifting away while stimuli remain visible, are advocate interpretations rather than experimentally established causal results.

The Reporting Summary documents open data/code, separate optimization and confirmatory samples, exclusions, modality-specific preprocessing, MRI quality control and multiple-comparison procedures. Its total 237 included datasets covers optimization plus confirmatory data and complete iEEG datasets; it must not replace the analysis-specific sample sizes. No diffusion-MRI connectome or microscopic connectivity map was measured.

A few internal reporting issues warrant restraint with fine-grained RSA claims. Supplement Table 16 includes a PFC letter-orientation GNWT-template correlation despite nearby prose describing none; its direct model-preference test is nonsignificant. Two PFC IIT-template entries appear interchanged between Tables 15 and 18. These inconsistencies were not silently reconciled or used to strengthen a theory verdict. The main conclusions above rely on the consistent comparisons and preserve the separation between template correlation and direct model preference.

Interpretation and limits

The cleanest outcome is a constrained challenge to two sets of biological commitments. The result does not establish a winner, falsify every possible IIT/GNWT formulation, or settle whether consciousness requires a particular physical substrate. The article expressly allows a theory's biological implementation to change while its mathematical/computational core remains.

The large sample, multiple modalities, task-relevance manipulation and explicit prospective criteria make these results more informative than a single favorable correlate. At the same time, the nulls remain conditional on region sampling, temporal windows, frequency bands, measurement sensitivity and the selected observable. Phase coherence, amplitude dependence, anatomical connectivity and IIT's intrinsic causal integration are different objects. The label “connectivity” must not collapse them.

The study did not record individual neurons, intervene on regions, compare synthetic brains with biological brains, assess selfhood, or track an individual's continuity through copying. Its individual-identity stimulus manipulation concerns which image was seen, not personal identity.

Relevance to consciousness, connectomics and Kurisutina

  1. A connection map is not the tested conscious state. The study's commitments concern content-dependent activity and interaction through time. Anatomical wiring can constrain these processes, but this paper supplies no proof that reconstructing anatomy automatically reproduces them.
  2. Observable information and causal necessity must be separated. Successful decoding shows that an external observer can recover some stimulus information from a signal. Failed incremental decoding does not show that the added region is unnecessary for producing experience.
  3. A replica needs claims matched to evidence. Reproducing reports, categories or autobiographical answers is behavioral evidence. This paper offers no bridge from those capabilities to subjective experience or preservation of the original person's identity.
  4. Methodological lesson: Before testing a consciousness-oriented model, specify which content, duration, region/system, observable, intervention and null result would count against it. Keep preregistered tests separate from informative exploratory repairs.

These are research implications, not findings that a connectome emulator or language-model persona is conscious. No source-person continuity claim follows.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.