Kurisutina

HippoRAG 2 and the meaning of artificial memory

Citation and version. Bernal Jiménez Gutiérrez, Yiheng Shu, Weijian Qi, Sizhe Zhou, and Yu Su. From RAG to Memory: Non-Parametric Continual Learning for Large Language Models. ICML 2025, PMLR 267:21497–21515. Official proceedings record, published PDF. The related preprint is arXiv:2502.14802; the retained and read version is the conference paper, not an assumed latest preprint or code release.

Reading scope. Complete main text and Appendices A–G, including prompts, ablations, example failures, implementation, and resource tables. Main-text figures/captions read; Figure 3 and Tables 4–5 and 12–14 visually checked against rendered PDF pages. References were not all read, and cited studies, implementation, data, and outputs were not independently audited or rerun. Source manifest.

Concise source findings. HippoRAG 2 uses LLM-extracted triples, linked phrase and passage nodes, embedding retrieval, an LLM relevance filter, and Personalized PageRank to select passages for a question-answering reader. Seven document benchmarks test factual, multihop, and discourse questions. With the principal reader, reported aggregate answer F1 is 59.8 versus 57.0 for NV-Embed-v2 retrieval. Supporting-passage recall@5 improves across three multihop datasets, with unequal gains. Ablations favor query-to-triple linking and passage nodes; filtering has a smaller average contribution and does not help every dataset. A corpus-expansion experiment adds documents while evaluating questions from one fixed segment; relative advantages persist while both methods decline on the harder multihop task. The graph and filtering terminology is motivated by human memory, but the outcomes are document retrieval and question answering. There is no human memory extraction, prospective intention execution, personal longitudinal adaptation, or biological mechanism validation in these experiments.

What the method actually transfers. The indexed source is already supplied text. The graph helps locate source passages, and the final reader still receives the top five passages. Thus a useful engineering hypothesis is that relations and shared entities improve navigation through known records. Successful retrieval does not show how unknown autobiographical content can be obtained from a person. Nor does it establish that a graph is sufficient to preserve their source judgments, accessibility, affect, or learning dynamics.

The relevance filter selects potentially useful triples. Its label, recognition memory, does not make it a verifier of historical truth or personal source attribution. If an input says that someone falsely remembers an event, flattening that sentence into an event assertion would lose a distinction this research program needs. Whether a particular implementation makes that error requires a direct test; it is not a failure demonstrated by the present benchmark.

Quantitative and methodological audit.

  • Different metrics. Figure 1's associativity F1 values, 62.96 versus 58.81, imply approximately 7.06% relative improvement but 4.15 absolute points. The abstract's 7% and introduction's seven-point wording should therefore not be treated as interchangeable. Separately, our arithmetic from Table 3 gives an average multihop recall@5 gain of (5.0 + 13.9 + 1.8)/3 = 6.9 percentage points across its three multihop retrieval datasets. This is a different endpoint and subset. The overall F1 gain from rounded Table 2 averages is 2.8 points. Weighting the seven displayed dataset scores by their question counts reproduces the reported averages after rounding; an unweighted dataset mean does not. These calculations are ours; none measures human memory.
  • Scope of the ablations. In Table 4, removing the filter lowers mean multihop passage recall from 87.1 to 86.4, whereas removing passage nodes gives 81.0. Changing query-to-triple to named-entity linking gives 74.6. These comparisons are conditional on the remaining pipeline, with interacting changes and no complete factorial design. They cannot rank universally necessary components or establish that a biological counterpart causes the improvement.
  • Splits and tuning. Appendix G reports hyperparameter tuning on 100 MuSiQue training examples, including comparison-method QA tuning. Table 5 separately reports selecting the passage reset weight using 1,000-query NQ and MuSiQue development sets. The relationship between those development sets and the main evaluated subsets is not fully resolved by the paper alone. The filter prompt is optimized using DSPy/MIPROv2; exact tuning examples and overlap need a code/data audit. This is a reproducibility question, not evidence that contamination occurred.
  • Expansion is one form of updating. Section 6.3 changes the retrieval corpus, not a participant's beliefs or the truth status of an earlier claim. It does not test explicit correction, cancellation, source contradiction, selective forgetting, or behavior triggered by an intention at a later opportunity. Robustness to extra distractors is valuable but addresses only part of a personal-memory system's requirements.
  • Failure denominators. Appendix E examines 100 cases selected for imperfect passage recall. Its component error percentages are conditional on that selected set; they are not whole-benchmark error rates. Phrases near the linked nodes need not imply the downstream ranking can recover the correct evidence.
  • Cost and comparison limits. The reported MuSiQue setup deploys its 70B model on four H100 GPUs. Resource tables omit model-weight memory. HippoRAG 2's reported indexing and query costs exceed dense retrieval; comparisons with other graph systems are conditional on the specified prompts, modes, retrievers, and hardware. These historical measurements do not provide a current deployment price or a universal performance ranking.
  • Printed inconsistencies. Independent visual review confirms that Table 11's second question concerns Philippe's grandmother while its candidate and filtered triples concern Bank of America/FleetBoston. It cannot substantiate the accompanying claim of correct phrase identification followed by a graph-search failure. In Table 12, LightRAG's 13.3 seconds divided by HippoRAG 2's 1.2 seconds is approximately 1108.3%, not the printed 1008.3%. These example/table problems warrant reconciliation, without independently invalidating aggregate benchmark scores. Smaller percentage mismatches may involve rounded inputs.

Proposed Amadeus comparison. Give a dense retriever and a graph retriever identical dated, attributed episode/report records and the same reader. Keep unknown event truth evaluator-only when testing extraction. Match source access and report both quality and acquisition/inference costs. Include arbitrary bindings, associative paths, contradictory reports, and explicit revisions. Test graph retrieval with and without source-preserving relation labels, rather than assuming semantic triples already distinguish an event from a belief about it.

Then separately test action: create an intention in an earlier interaction, deliver a relevant environmental cue amid an ongoing task, and measure whether the recipient performs the action within its valid window. All systems must receive the same cue stream and execution opportunities. Prompting the model at the opportunity with a question that restates the intention would instead test prompted recall. A scheduler or persistent state can supply needed functions without copying hippocampal anatomy; compare such simpler alternatives directly.

Finally, evaluate fidelity to the person's observed changes separately from improvements in objective record accuracy. A system that never forgets supplied passages may be an excellent archive while predicting the person's later accessibility poorly. This distinction determines the benchmark, not the biological name assigned to a module.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.