Citation and version. Reese Kneeland, Cesar Kadir Torrico Villanueva, Tong Chen, Jordyn Ojeda, Shubh Khanna, Jonathan Xu, Paul S. Scotti, and Thomas Naselaris. MIRAGE: Robust multi-modal architectures translate fMRI-to-image models from vision to mental imagery. Peer-reviewed Methods article, PLOS Computational Biology 22(5):e1014263, published 22 May 2026. Published article. This note uses the published paper and its S1 Text, not the differently authored preprint. Article and illustrative assets are licensed CC BY 4.0; see the supplement's explanation of replaced source images.
Reading scope. Full main text, methods, all table/figure captions, and all 18 pages of S1 Text, including Appendices A.1–A.17. Visually checked main Figure 2, Figure 5, Tables 1–2, supplementary median/worst examples C–F, training-scale plot N, diffusion plots O–P, task example S, and the image-substitution disclosure. Other captions/tables were read from extracted text; not every plot was independently digitized or visually audited. Read the public repository README pinned to 7dbc13fb888df0845ea40c4a8c7375cf0bce626f; implementation, raw results, and datasets were not audited or rerun. Provenance.
Factual core. MIRAGE maps fMRI through ridge regressions into image, text, and low-level features that guide image generation and candidate selection. Its main imagery evaluation uses four NSD participants, individually trained decoders, and 18 learned letter-cued targets: six shapes, six complex images, and six concepts. Other participants' imagery data guide regularization and image-filter choices. Main training uses roughly 40 hours of visual data per person; test reconstruction averages repeated trials. In a human two-choice evaluation, overall imagery identification is 78.30%, versus 73.95% for Brain Diffuser and 56.96% for MindEye2. Numerical reconstruction metrics average generated outputs; selected illustrations show the best examples. Median and worst examples are also supplied. Different metrics and stimulus types favor different methods. These results support cross-decoding of instructed imagery within this benchmark, while leaving unknown autobiographical content and personal imagery detail beyond the known target untested.
Method audit and quantitative boundaries.
- People, sessions, and stimuli. Main evaluations use subjects 1, 2, 5, and 7; subjects 3, 4, 6, and 8 supply imagery-informed tuning. NSD-Imagery is a separate session from the visual-training data. This is a useful session distinction, and ridge weights are trained on vision. It is not an imagery-blind development process: imagery performance selects regularization and guides architecture development. The same small target set underlies tuning and evaluation; there is no demonstrated generalization to unrestricted novel mental content. The five COCO complex targets come from shared1000; exact training/test image membership was not independently audited here.
- Cue protocol. Participants learn a letter-to-target mapping, see the letter and empty frame, and imagine the corresponding target for three seconds. Vision and imagery occur in separate runs; recent vision runs refresh the targets. Separating runs addresses immediate visual-response carryover, but does not by itself isolate vivid imagery from other cue-triggered semantic or memory activity. A cue-only recipient supplied the learned mapping would already know the requested target; that baseline asks about incremental acquisition, rather than whether a visual-trained decoder responds to the brain signal.
- Repetition and calibration. The supplement reports eight repetitions for seen targets and sixteen for imagined targets, with a separate repetition-scaling analysis. These are repeated attempts at known targets, not sixteen independent unknown memories. The smaller-training-data curve is shown for subject 1; its three-hour relative advantage cannot establish a general minimum calibration requirement. Acquisition cost includes visual training, target familiarization, repeated imagery, scanning, and inference.
- Selection stages. Sixteen generated candidates are ranked using a brain-decoded image embedding; that automatic step does not consult the target image at inference. The main text reports averaging ten final outputs per test case, but A.6 also gives five-output MIRAGE wording alongside inconsistent table references; the precise per-table count remains unresolved without a reproduction audit. Ten is explicit for the human evaluation in A.14. A separate target-informed ranking selects best/median/worst illustrations. Keeping these stages distinct avoids either mistaking the pictures for unselected performance or wrongly treating automatic selection as oracle leakage.
- What identification means. Feature-based two-way comparisons use other benchmark reconstructions as distractors; their pool differs from shared1000 benchmarks. Human two-choice distractors are matched for source participant, method, stimulus type, and vision/imagery condition. Neither result is the fraction of pixels or memory details recovered. Same-type shape discrimination is a useful constraint beyond a generic natural-scene category; concept matching still lacks an image of what the person actually imagined.
- Statistical scope. The human study recruits 500 raters and excludes five for attention-check failures. Many ratings and generated outputs reuse the same four scanned people and small target set. The text reports small p-values, but does not clearly specify the inferential model or resampling units for all primary comparisons. Table A supplies standard errors rather than a complete test specification. Rater precision cannot be treated as evidence from hundreds of independent brains. The supplement also contains five-versus-ten output and table-reference wording that needs reconciliation before reproduction.
- Prior controls. Comparing reconstructions generated from different brain patterns while sharing the generator supports target-related signal. It does not establish that every depicted detail is signal-derived. The diffusion-strength experiment retains learned VDVAE/text/generator structure, removes image guidance, and varies one parameter within that altered pipeline. A nonsignificant imagery slope is not equivalence; strength values in different architectures do not measure a common fraction of prior dependence. The evidence does not establish prior-free reconstruction.
- Published pictures are not exact target records. Appendix A.17 explicitly says some complex-stimulus ground-truth panels were replaced by generated proxies for republication. The authors state that training and quantitative evaluation use the original images. Thus the published side-by-side panels cannot support exact detail-by-detail comparisons to the actual scanned stimulus. This is a documented illustration limitation, not evidence that the numerical experiment used proxies.
Interpretation for memory extraction — our reasoning. A result may show that brain responses distinguish target categories while leaving their idiosyncratic details unresolved. A generator can convert a small decoded distinction into a richly detailed output. That output is a hypothesis conditioned on a signal and learned priors; its visual specificity is not automatically acquired specificity. Repeated generated agreement can reflect the same shared prior rather than independent corroboration.
The shape results make it inappropriate to dismiss this method as merely generating an arbitrary natural scene. Conversely, matching a known target better than a distractor does not establish a faithful image of the participant's current recollection. Their remembered image may itself differ from the stimulus. An extraction study should retain three references where available: original event evidence, separately obtained personal reports, and consequences in subsequent behavior. None alone directly reveals every feature of subjective imagery.
Proposed discriminating test. Hold cue identity, category, public instructions, and decoder-accessible records fixed while varying an arbitrary, privately experienced detail. The detail must be absent from training, candidate descriptions, test wording, and all other recipient inputs. Evaluate its recovery on future episodes, using original records only for scoring and obtaining personal reports separately after decoding. Also test against same-category, same-cue alternatives that differ in that detail. A broad category-matched image is insufficient for this endpoint.
Compare a cue-and-prior recipient, a recipient with the acquired behavioral record, and that same recipient with the neural channel. Separately compare equal-cost acquisition policies, since extra behavioral elicitation can change the person's memory. Freeze generation and candidate-selection rules before evaluation, retain all outputs, and report uncertainty at the person/episode level appropriate to the claim. Add an autonomous action whose outcome depends on the recovered detail; score source-person fidelity separately from event accuracy. These proposed controls would test information acquisition and functional use beyond the published benchmark.