Kurisutina

Shared computational principles for language processing in humans and deep language models

Nature Neuroscience 25:369–380, doi 10.1038/s41593-022-01026-4; open access, PMC8904253. Read by the main session for research direction R3, 4 October 2026. Provenance: papers/amadeus/goldstein2022_shared_principles.provenance.json.

What was read

  • Read in full: the Europe PMC full-text XML, converted to text: abstract, main text, figure captions, every Extended Data caption (S1–S10), Methods, statements, and the reference list.
  • Not read: Supplementary Information (Appendix I, the decoder architecture; Appendix II, the word list), the Reporting Summary, and the figures as images.

What they did

  • Behaviour. 300 online participants, in six groups of 50, predicted every next word of a 30-minute podcast ("Monkey in the Middle", This American Life), 5,078 words, in a sliding 10-word window. Their predictions were compared with GPT-2's (the 48-layer version) and with n-gram models.
  • Brain. Electrocorticography (ECoG: electrodes on the cortex) in 9 epilepsy patients, 1,339 electrodes (1,106 left hemisphere), while they listened freely to the same podcast with no instruction to predict. Signal: 70–200 Hz broadband power, 512 Hz.
  • Encoding and decoding.
    • Encoding: linear models predict each electrode's response to each word from embeddings (arbitrary, GloVe or GPT-2 contextual), at lags from −2 s to +2 s around word onset. 10-fold cross-validation; phase-randomisation null; FDR q < 0.01.
    • Decoding: a neural network maps the activity of several electrodes back into embedding space, and word identity is classified by cosine distance (ROC-AUC).

Main results (verified)

  • Human and model next-word predictions agree.
    • Mean human predictability was 28%; about 600 words exceeded 70%.
    • Human and GPT-2 predictability correlate at r = 0.79, rising from 0.46 with 2 words of context to an asymptote at about 100 words. Top-1 predictions match 49.1% of the time.
    • The calibration curves are similar: both are under-confident and above 95% correct when the stated probability exceeds 40%.
  • The brain carries next-word information before the word is heard.
    • Encoding of upcoming-word identity was above chance up to 800 ms before onset with GloVe; robust from −1,000 ms in the figure. The peak was 150–200 ms after onset.
    • It held with arbitrary embeddings, after removing repeated bigrams, and beyond the previous word's embedding.
    • Preprocessing leaks at most about 93 ms from the future, verified on an impulse response. That is far shorter than the pre-onset effects.
  • Wrong predictions are visible as what the listener expected. Before onset, activity encoded the predicted word even when the prediction was wrong. After onset, it encoded the word actually heard. "Correct" was defined by GPT-2's top 5 (62% of words); top-1 and human-prediction splits are in Extended Data Fig. 8.
  • Confidence before, surprise after. GPT-2's entropy (uncertainty) correlates negatively with activity before onset. Its cross-entropy (surprise) correlates with activity after onset, peaking around 400 ms. These are partial correlations, each controlling for the other.
  • Contextual embeddings fit better than static ones.
    • GPT-2 embeddings gave significant encoding in 208 left-hemisphere electrodes, against 160 for GloVe. The 71 electrodes added were in higher-order language areas (IFG, temporal pole, pSTG, angular gyrus) and motor cortex.
    • Averaging a word's contextual embeddings across its occurrences drops performance to GloVe level. Shuffling them between contexts also reduces it.
    • Decoding word identity: AUC 0.74 with GPT-2 against 0.68 with GloVe or arbitrary embeddings, peaking at +150 ms. Above chance from about −1 s; at chance beyond 2 s.

Limits

  • Nine patients with clinically placed electrodes (coverage set by clinical need; left-hemisphere bias). Electrodes are pooled across people, so nothing here separates what is individual from what is shared.
  • One 30-minute story, passive listening. The "predictions" are GPT-2's or the online group's, not each patient's own reported predictions.
  • Inconsistent electrode bookkeeping: the text says "164 of 1,065 electrodes removed" in one place and "66 electrodes … removed" in another, against 1,339 implanted. The printed entropy formula lacks its minus sign. Neither affects the main claims, but both are sloppy.
  • The Conclusion's view that language models "cannot think … they simply echo the statistics of their input" is the authors' opinion in 2022, not a result.
  • Later critiques of pre-onset encoding were not read here.

What it means for Amadeus (inference)

  • A language model's internal space is a usable coordinate system for live brain activity. That is the bridge E3 needs: neural recordings during speech or thought can be tied to meaning by linear maps into LM embedding space, and decoded back. That is "tie the self-supervised neural model to meaning through anchors" in its simplest working form.
  • The brain shows the person's expectations, including wrong ones. Pre-onset activity carried what the listener expected, not what came. A person's predictions about the world are exactly what a person slot must hold; a recorder can see them before they are spoken. Caveat: here the "person" was a group proxy. A per-person version needs the person's own predictions as labels.
  • No individual specificity was tested. The method pools electrodes across nine people. Whether this person's neural encoding differs usefully from a shared map is the question Charest 2014 asks for vision. For language, it is open in what has been read so far.
  • Invasive recording. ECoG is surgical. It exists only in epilepsy monitoring, typically days to weeks of implantation. The idea of a month under electrodes has precedent in clinical monitoring (R1's question), but not for healthy volunteers.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.