Kurisutina

Prompted imagery and the limits of reconstruction evidence

Research date: 2026-09-22.

Citation and versions. Hugo Caselles-Dupré, Charles Mellerio, Paul Hérent, Alizée Lopez-Persem, Benoit Béranger, Mathieu Soularue, Pierre Fautrel, Gauthier Vernier, and Matthieu Cord. Mind-to-Image: Projecting Visual Mental Imagination of the Brain from fMRI. The ten-page workshop paper identifies acceptance at the ICML 2024 AI for Science workshop, not the main conference. The retained PDF is the distinct, twelve-page arXiv v5, dated 28 May 2024; arXiv describes it as a work in progress. DOI identifies the preprint. Its verified license is CC BY 4.0.

Acquisition and reading. Read all main text, methods, results, captions, references, and appendices in both versions. Visually inspected all five figures and Table 1 in v5. The workshop PDF was fully readable through browser extraction, but shell download returned HTTP 403 and browser screenshots failed; its 615 extracted lines are retained, not a byte-identical PDF. No workshop figure pixels were independently inspected. No raw fMRI, code, trial-level outputs, or external supplementary materials were obtained. Local v5 PDF, v5 text, complete workshop text, and provenance.

Source findings (approximately 150-word factual core). The protocol uses 1,200 surrealist portraits and landscapes and approximately six hours of scanning. For weak imagery, recently viewed images are flashed again for 100 ms before five seconds of imagined recall; 75 image–brain pairs are reserved for validation/evaluation. A modified MindEye decoder maps fMRI to diffusion-model inputs. Its best reported weak-imagery model produces images classified into the correct portrait/landscape category 91% of the time; CLIP two-way identification is 68.5%. Freezing this decoder and applying it to instructed novel imagery gives 88% category classification. Strong-imagery prompts combine portrait/landscape with emotions; oral descriptions supply qualitative comparison. Only brain data reportedly enters the reconstruction model at inference. Detailed strong-imagery accuracy lacks an external image target and systematic quantitative scoring. The paper acknowledges variable detail fidelity. ArXiv v5 additionally reports an inconclusive MindEye2 fine-tuning attempt, with CLIP two-way identification of 47.8%. These results concern elicited current imagery, not retrieval of unknown autobiographical events. Full v5, §§3–5.

Version audit. Compared with the accessed workshop snapshot, v5 adds §4.1.2 on fine-tuning a multi-subject MindEye2 model, the associated fourth experimental row in Table 1, a discussion of disagreement between image metrics, and Appendix B's training details. It renames masks A, B, and their union as Vision, Imagination, and their union. The shared category accuracies and original three experimental rows agree. Appendix B specifies one A100, 500 epochs, batch size 16, and 15–25 hours per training run. The versions are not interchangeable; the MindEye2 findings and these training details are cited to v5. Their chronological ordering is not inferred from page count or workshop status.

Methodological audit — source facts and what they leave unresolved.

  • Participant and trial accounting. Methods explicitly describe a mask for one subject, and discussion refers to the subject's repeated imagery becoming clearer. Elsewhere generic plural wording appears; there is no formal participant count/demographic table or across-person replication. Strong imagery crosses two categories with ten emotions, each imagined for six seconds, with ten cycles. Actual retained trial/output denominators, exclusions, repeated-trial averaging, and independence of evaluation units are not specified. Do not turn this description into 200 verified independent tests. The text also describes six hours both for weak imagery and for the combined protocol, without a complete scan allocation.
  • Selection and leakage boundaries. The 75 weak-imagery pairs serve validation and evaluation; models/masks are compared there, without a separately described final weak-imagery test. Manual masks use GLM event-versus-rest responses, but training-only mask construction and run/session separation are not documented. These are unresolved dependencies, not proof of label leakage. Strong-imagery inference is reported to use a frozen model and no strong-imagery training targets. The paper does not specify enough stimulus/run accounting to verify item, run, and preprocessing independence throughout.
  • What the metrics identify. The 91% and 88% values come from a fine-tuned ResNet50 classifying generated images as portraits or landscapes; they are not percentages of detailed mental images correctly recovered. Two-way feature comparisons in Table 1 evaluate weak imagery, with no reported within-category distractor analysis. Nominally above-50% results are not accompanied by reported confidence intervals or an inferential test. Comparisons with original MindEye on NSD involve different tasks, data, and image distributions; the authors explicitly disallow a direct performance comparison. Pixel correlation, structural similarity, and feature distances favor different models, so no single ordering establishes overall fidelity.
  • Cue and report channels. Weak imagery follows both recent full image exposure and a brief visual cue. Separate GLM event estimates do not, by themselves, demonstrate clean separation of cue/perceptual and imagery contributions; the necessary design diagnostics are not supplied. Strong imagery is category/emotion instructed. Prompt presentation is described as verbal in methods and on-screen/written in captions/conclusion. Oral descriptions are collected during or after scanning and used for evaluation, not reported as decoder inputs. Description timing relative to viewing reconstructions, blinding, rater counts, scoring rules, and competing-image comparisons are not reported.
  • Visual evidence. Figures 3 and 4 are explicitly curated examples. Figure 4 shows four generated images with category/emotion prompts, without verbatim oral descriptions or independently scored details. Appendix A displays 15 uncurated weak-imagery pairs, not all 75 validation items; there is no corresponding uncurated strong-imagery panel. Visual plausibility therefore does not estimate average recovery of private detail. All these observations were checked against v5 figure pixels.

Interpretation for extraction — our reasoning. Three targets must be separated: the instructed category, the participant's current within-category content, and the historical truth of an event. The experiment provides its clearest evidence about the first. A decoder receiving only fMRI can still convey information already available to an observer from the task instruction. Neural input is not equivalent to incremental private information beyond that instruction, and a diffusion prior can supply vivid unsupported details.

Above-chance pairwise identification alone does not resolve this problem. As an illustrative null, suppose the distractor category is sampled independently of the target category and matches it with probability one half. A perfect category decoder with no within-category discrimination then achieves 75% expected two-way accuracy: it wins against the other category and guesses within its own. This population/category-sampling assumption is not identical to sampling a distractor uniformly from a finite balanced item pool after excluding the target; with ten items per category, the latter gives approximately 76.3%. Neither number estimates the paper's actual null, whose distractor construction and validation balance are insufficiently specified. The example shows why its 68.5% CLIP result needs a within-category comparison before being interpreted as item-specific recovery.

The study tests construction of new imagery under instructions and immediate cued recall of presented images. It does not establish decoding of spontaneous dreams, access to an unknown autobiographical archive, or verification that a remembered event happened. Successful decoding of current imagery would still leave event truth unresolved.

Falsifiable acquisition and transfer test — proposal. Hold the public category/emotion prompt constant while the participant privately chooses unpredictable details and commits an independent record before seeing reconstructions. Keep that record and oral descriptions outside decoder inputs; their agreement is a criterion for reported content, not privileged access to imagery. Freeze masks and model selection before held-out runs and episodes. Compare equal-budget generation from the public prompt alone, actual neural observations, and neural observations shuffled across different episodes sharing the same prompt. Also test instruction reading without imagery, with appropriately separated timing, to isolate the additional contribution of imagining.

Precommit scoring and any sample-selection rule; evaluate every designated output with blinded, within-prompt alternatives. Report single-trial and repeated-trial-averaged results separately, with participant/event-level uncertainty. Improvement on unpredictable details over these controls would support content acquisition beyond the public prompt. Failure would leave category decoding intact while weakening the richer extraction claim.

Finally, give a recipient system only the acquired record and test choices that depend on those held-out details. Score information accuracy, prediction of the source person's later report/choice, and the recipient's own use separately. Compare the recipient with the source person's matched continuation behavior. Generic task success without additional personal information would not show functional transfer of the source's memory; matching both independently verified details and their personal consequences would be stronger, still task-bounded evidence.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.