eLife 9:e56601, doi 10.7554/eLife.56601; open access (CC BY), PMC7266639. A review ("Perspective") of the authors' own method and results. Read by researcher R4e (innate against learned) for research batch R4, 5 October 2026. Provenance: papers/base/haxby2020_hyperalignment.provenance.json.
What was read
- Read in full: every line of the PMC XML, converted to text: abstract, impact statement, all sections, the two equations, Table 1, figure captions, funding and the reference list.
- Not read: figures as images. The paper has no supplement.
- Nature of the numbers: this is a review. Nearly every number below is the authors' report of their earlier papers (Haxby 2011, Guntupalli 2016 and 2018, Feilong 2018 and 2019, Nastase 2017, Jiahui 2019, Van Uden 2018). The cited source is named with each number; none was read here.
The idea
- Shared information, idiosyncratic layout. The same information is encoded in every brain, but in fine-scale cortical patterns whose layout differs between people. Aligning brains by anatomy misses it.
- Hyperalignment builds a common, high-dimensional information space. Each person gets their own transformation matrix R_i that maps their cortical loci (voxels or vertices) into the common space's dimensions. Equation 1: the common model M is the mean of the transformed individual data, with each R_i chosen to minimise the distance between B_i R_i and M.
- "Information" is vector geometry: the pattern of similarities and distances among responses to different stimuli or time points (or among connectivity patterns). The main algorithm (generalised Procrustes analysis) uses rotations, which preserve each person's geometry exactly; the shared geometry is what survives.
- Topography as mixtures: each column of R_i is a person-specific topographic basis function. A cortical pattern in one person is the common-space pattern times R_i transposed (Equation 2). The same common-space pattern predicts a different map in each brain.
- The authors' central claim: "The fundamental property of brain function that is preserved across brains is information content, rather than the functional properties of local features that support that content." Unit tuning is shaped to support the shared geometry, not the other way round.
Variants
- Searchlight whole-cortex hyperalignment: local models in overlapping patches, aggregated, so information is only remixed between nearby locations.
- Dimensionality reduction (PCA; the shared response model, SRM, of Chen 2015, which fixes a low dimension in advance) can keep or improve between-subject classification by discarding noise.
- Regularised CCA and joint SVD also work; slight relaxation of orthogonality helps.
- Connectivity hyperalignment: aligns patterns of functional connectivity instead of responses. It needs no shared stimulus, so it works on resting-state data and across different experiments.
What the data need
- Rich, naturalistic input. A common model estimated from responses to a full-length movie generalises to other experiments; one estimated from a controlled experiment with few stimuli does not (Haxby 2011).
- Between-subject classification of dynamic movie segments in ventral temporal cortex was twice that for unpredictable static images (Haxby 2011). The authors attribute this to the variety of brain states movies evoke (perception, action, narrative, speech, memory, attention, semantics, social cognition) and to the predictions they build.
- Whether a model built from resting-state data generalises to many tasks is not yet established.
Main results reported (as cited)
| Measure | Anatomical alignment | Hyperalignment | Source cited |
|---|---|---|---|
| Classifying movie time segments across subjects (chance < 0.1%) | 1.0% | 13.7% (response); 10.4% (connectivity) | Guntupalli 2018 (the only source cited in that paragraph; the numbers carry no citation of their own) |
| In small regions | Very low | More than an order of magnitude higher | Guntupalli 2016, 2018 |
| Intersubject correlation of movie time series | — | Nearly doubled | Guntupalli 2016 |
| Intersubject correlation of local representational geometry | 0.21 | 0.32 (response); 0.31 (connectivity); mean increase 46–75% in other studies | 0.21–0.32: as above; 46–75%: Guntupalli 2016, Van Uden 2018 |
| Intersubject correlation of connectivity profiles | 0.15 | 0.58 (response); 0.67 (connectivity) | As above |
| Classifying still images of faces, objects and animals in a movie-derived common space | — | Equal to or better than within-subject classifiers | Haxby 2011; Guntupalli 2016 |
| A person's face-selective map, predicted from other people's data | — | Local correlation with their own map above 0.80 in posterior ventral temporal, lateral occipital and posterior superior temporal cortex | Jiahui 2019 |
- Granularity: the common model keeps distinct profiles at about one voxel (about 3 mm) in occipital, temporal, parietal and prefrontal cortex. For connectivity, the between-subject point-spread after hyperalignment equalled the within-subject one across independent sessions (Guntupalli 2018).
- Classic categories are a small part of the space. In ventral temporal cortex, about 35 principal components captured the information in one hour of movie (Haxby 2011). All 35 together explained 58% of movie-response variance (cross-validated). The face-against-object dimension explained 7% (12% of the model's share); the first component alone 22%. An animacy dimension explains over 50% of variance for still images of animals and objects, but 6% for the movie.
- Individual differences that survive alignment are real.
- After hyperalignment people are much more alike, but their residual differences are more reliable across independent data than differences in anatomically aligned data. This held in 82% of searchlights, and it comes from fine-scale patterns (Feilong 2018).
- These fine-scale differences were more predictive of general intelligence than coarse-scale ones (Feilong 2019, a conference abstract).
- The authors' conclusion: "the most informative individual differences in cortical function reside in fine-scale patterns."
Interpretation the authors offer
- Degeneracy: structurally different elements perform the same function. It is typical of systems that evolve, develop or learn; designed systems avoid it.
- Deep against surface structure, as in language: the same meaning in different words and languages, the same information in different topographies.
- Deep networks show the same thing: networks trained from different initialisations or architectures have different individual units but share the similarity structure of their layers (citing Kornblith 2019).
- How experience shapes unit tuning to support the shared geometry is open; topographies may arise with development and experience.
Limits
- A review of one group's method. Effect sizes come from the cited papers, mostly small fMRI samples of young adults watching Hollywood movies; sample sizes are not given here.
- All of it is "shared after growing up". No developmental, cross-cultural or genetic data. The paper cannot say whether the common space is innate or learned from common experience.
- The intelligence result is a conference abstract.
What it means for Amadeus (inference)
- The base-and-slot split has a measured precedent. Hyperalignment factors every brain into a shared information space (the base) plus a person-specific transformation (part of the slot). This is the shape of the Amadeus design, and the review shows it working across the whole cortex.
- Two kinds of individuality, and the replica needs only one. (a) Where the shared information sits in a given brain: the transformation R_i. For a functional replica with no cortex this is irrelevant: it is the "neural fingerprint" report A already set aside. (b) What information differs after alignment: smaller, but reliable and predictive of cognition. That residual is the slot's real content.
- Train the shared part on rich, everyday input. Common models built from naturalistic, varied, predictable input generalise; models from narrow controlled tasks do not. For the base: a broad diet of what every person meets, not narrow task data.
- The shared space is high-dimensional, and textbook categories are a small fraction of it under natural conditions. The vocabulary engine should not be reduced to a few hand-chosen dimensions.
- Different implementations, same geometry. If networks trained on the same data share their geometry whatever their initialisation, then the base is defined by the geometry it reaches, not by its weights. This supports P7: define the base by a battery, not by its architecture.