Kurisutina

Semantic reconstruction, completed source audit

Citation: Jerry Tang, Amanda LeBel, Shailee Jain, and Alexander G. Huth. Semantic reconstruction of continuous language from non-invasive brain recordings. Nature Neuroscience 26, 858–866. Peer-reviewed article; version of record published 1 May 2023. Publisher, DOI: 10.1038/s41593-023-01304-9; public accepted manuscript, PMCID PMC11304553, NIHMS2005151.

Reading and version scope: Read the fresh PMC manuscript's complete main text and Methods, all Figures 1–4 and Extended Data Figures 1–10 legends, and Table 1. Read the complete 36-page published supplement, including Supplementary Figure 1 and every row/caption of Tables 1–7, plus the complete four-page reporting summary. Visually inspected main Figure 4, Extended Data Figure 7, Supplementary Figure 1, and selected table pages to check highlighting and answer keys. The current publisher's Table 1 values/caption and Figure 3 caption agree with the manuscript on the audited points. This is not a complete comparison against a version-of-record main PDF: the accessible main text is the accepted manuscript, whereas supplements and the two separately accessible figure/table pages came directly from the publisher. References were not exhaustively reread, and the supplementary video, code, data, and individual statistical computations were not inspected. Provenance records all 14 retained source/extraction artifacts, exact URLs, hashes, and limits. Public access does not establish an open redistribution license. Original papers/04* and the original summary are preserved.

Evidence, concise paraphrase: Three intensively sampled adults provided roughly sixteen hours of story-listening fMRI for personalized models. A language-model prior proposed continuous text, and predicted brain responses ranked candidate continuations. Generated text scored above the reported semantic null for heard stories, cued retellings of supplied narratives, and silent movies. Imagined stories were one-minute Modern Love excerpts, not the participants' own life events; five-way identification of their retellings was perfect, without implying perfect reconstruction. Perceived-story word error rates were 0.924–0.941, and supplementary outputs contain substantial factual substitutions. Null sequences retained brain-predicted word timing. Above-chance story decoding occurred with one training session, approximately one hour. Coverage and identification measures largely flattened around 7.5 hours; overall similarity continued improving. Development included qualitative examination of one participant's principal test story before the pipeline was frozen for other participants, stimuli, and experiments. These results establish constrained semantic reconstruction, without establishing extraction of previously unreported autobiographical details or incremental benefit over a person's report.

Method audit and interpretation:

  • What was imagined: Methods specify five Modern Love segments excluded from model training. Participants learned segment IDs, heard an ID cue, had ten seconds to prepare, and imagined telling its segment from memory for one minute; every segment occurred twice in one scan. Rehearsal dose, a memorization criterion, and the exact timing/order of spoken reference collection are not reported. Quantitative references were separately recorded retellings by each participant. Supplementary Table 2 instead displays the instructed source segments and reports similarity to both sources and retellings; it prints the second repeats only. Thus “cued retelling from memory” is supported; spontaneous autobiographical recovery is not. Narratives can be autobiographical for their original storytellers without being autobiographical memories of scanner participants.
  • Generation versus identification: Words were generated by the language/encoding pipeline before comparison against the five reference transcripts. This is not simply a five-template selector, but its 100% post hoc identification result remains a five-way test over ten trials per participant. The cue itself identifies the experimenter's source story, so identification alone cannot measure content added beyond cue knowledge. Information uniquely preserved in each participant's retelling is a more demanding target, not an outcome established by that identification score.
  • Meaning scores versus accurate details: BERTScore used its recall variant, with training-story IDF weighting, over overlapping twenty-second windows centered every second. Reported percentages of significant time points therefore do not mean percentages of correct words, propositions, or independent episodes. Table 1 gives BERTScore 0.7899 for the null versus 0.8077/0.8104/0.8116 for S1/S2/S3; corresponding WER is 0.9637 versus 0.9407/0.9354/0.9243. The translation comparison is a reference procedure, not a bound on recoverable information. Supplementary outputs show that above-null similarity can coexist with invented people, actions, and circumstances. The reader experiment found above-chance answers to nine of sixteen questions using S3's decoded text; it does not provide a general factual-accuracy rate or a neural-versus-self-report comparison.
  • What the null controls: Methods explicitly preserve brain-predicted word times, the language prior, and the continuation procedure while replacing neural likelihood rankings with random values. Consequently, Table 1's shorthand description of using no brain data is incomplete. Methods also specify a null beam of ten sequences versus the actual decoder's beam of two hundred, despite saying the ranking is the only difference. As described, the comparison jointly changes candidate ranking and search budget. Without a matched-beam comparison, the full score increment cannot be attributed specifically to brain-based ranking. This unresolved implementation-description mismatch does not show that the result disappears. The null also retains timing information from the brain and does not exhaust cue-conditioned alternatives.
  • Development exposure: Parameter search used S3's separate calibration story, From Boyhood to Fatherhood. Methods and the reporting summary also disclose qualitative exploratory analyses on S3's Where There's Smoke, which is the principal quantitative perceived-speech test story. The pipeline was frozen before viewing other participants, stimuli, and experiments. This supports subsequent generalization tests but prevents labeling every reported test result untouched by development. It does not show numerical optimization on test labels or contamination of all later tests.
  • Training burden and curve shape: Supplementary Table 5 explicitly uses one hour from one of fifteen training sessions, and the main text reports above-chance performance at that dose. Extended Data Figure 7's roughly 7.5-hour plateau concerns significant-window coverage and identification percentile rank. Figure 4b depicts continuing, diminishing returns in story-similarity scores. None defines an application-specific useful-performance threshold or a minimum required calibration time. Supplementary Figure 1 checks overlapping versus nonoverlapping training subsets using simulated responses from the fitted encoding/noise models; it addresses that sampling explanation within the model, not all causes of saturation. These learning curves do not prove a biological information ceiling.

Original scientific reasoning: The pipeline generates language related to target content, and task contrasts such as following the attended speaker support a role for measured activity. However, the reported null does not isolate the size of the contribution from brain-based candidate ranking: its search budget also differs. The unresolved extraction question is how much verified, event-specific information unavailable from the cue or existing report neural measurement supplies under matched conditions. The study's targets and metrics do not identify that quantity. Its three main participants were willing lab volunteers with intensive calibration; anatomically aligned cross-person failure is specific to that approach, and smoothing fMRI is not a measurement with a portable optical device. Neither finding supports an unconditional deployment limit or a demonstrated low-burden alternative.

Proposed extraction controls, not reported experiments: Freeze model and evaluation choices before new people and episodes. Give multiple distinct episodes the same nonspecific cue; hold out entire episodes, not adjacent windows, and separately score source, people, actions, temporal order, and unsupported additions. Compare the same language prior under matched beam/search budgets with (i) cue and report alone, (ii) cue plus predicted timing, (iii) content-shuffled neural signals, and (iv) the intact signal. Preserve timing, acquisition noise, and other nuisance structure where the contrast requires them. Grade blind against independently recorded event details, including details omitted from an initial report, then compare added verified information against equal time spent obtaining another behavioral report. Report dose–utility curves with uncertainty across people and episodes. Above-chance five-way identification, smoother prose, or higher semantic similarity alone would not meet that extraction criterion.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.