Citation: Di Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang, Kai-Wei Chang, and Dong Yu. LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory. Peer-reviewed ICLR 2025 conference paper; originally released as arXiv:2410.10813 in 2024. Official proceedings record, official PDF, repository.
Reading scope: Read the entire 28-page proceedings PDF, including main text, methods, ethics/reproducibility statements, references, and Appendices A–E with all prompts, tables, and figure captions. Visually checked Tables 3, 6, 9, and 10 against the PDF. Read the current repository README and verified its bytes against commit 9e0b455f4ef0e2ab8f2e582289761153549043fc (11 May 2026). No benchmark data, implementation code, model outputs, cleanup change log, or V2 paper/data were audited; no experiments were run. Provenance records retained files and hashes. Downloads used verified HTTPS. An article-specific redistribution license was not independently established.
Evidence, concise paraphrase: LongMemEval provides 500 questions about information extraction, cross-session synthesis, temporal reasoning, knowledge updates, and abstention. LLM-generated backgrounds seed questions curated by experts; simulated evidence conversations are inspected and approximately 70% edited. Histories mix evidence with ShareGPT, UltraChat, and additional simulated sessions. Standard settings contain roughly 115,000 tokens or 500 sessions. A GPT-4o answer judge is compared with human judgments on 30 questions per type using outputs from two models. Experiments separate indexing, retrieval, and answer generation. Retaining original conversation values while augmenting retrieval keys with extracted facts often improves performance. In Table 3, GPT-4o with ten retrieved rounds rises from 0.670 to 0.720 accuracy. Correct retrieval still sometimes produces wrong answers; oracle evidence improves but does not perfect performance. These are controlled evaluations of textual conversational information, rather than observations of real people's longitudinal memory or demonstrations of neural extraction. Paper-era commercial-system tests occurred in August 2024.
Method audit and interpretive limits:
- What is human: Expert involvement curates questions, evidence, timestamps, and phrasing. The evidence histories do not constitute prospectively collected personal lives. The nominal filler mixture is 25% ShareGPT, 25% UltraChat, and 50% simulation; public conversational data in the mixture do not make an entire constructed history one real person's record. Attributes and supplied answers shape the construction before evaluation.
- What the targets establish: Updated-answer correctness can be achieved by selecting the newest statement during reading; it does not require erasing an old record or changing model weights. Thirty false-premise questions test a bounded form of abstention. This does not establish calibrated confidence across arbitrary unknown, ambiguous, or genuinely forgotten experiences. Personalized-answer rubrics test use of supplied user information, not identity continuity or human choice prediction.
- Judge validity: Table 6 contains 206/210 and 204/210 agreements with human judgments, totaling 410/420 (97.6%, calculated from its cells). Two individual task/model cells are 27/30. This is finite, task-specific agreement, not universal judge reliability. The paper does not describe a separate held-out split for judge-prompt development. Its scoring rules tolerate small date offsets and some inclusion of old information alongside the required update; strict factual precision would need a distinct metric. The judge receives the reference answer and model response, rather than independently reconstructing evidence from the full history.
- Development versus evaluation: Prompts were manually tuned after inspecting examples; extraction demonstrations include simulated sessions and ShareGPT. No independent development/test partition for the design search or pretraining-contamination audit is described. This is a limit on claims of untouched-test generalization, not evidence that answer leakage or contamination occurred. The benchmark was publicly released in 2024, so present-day evaluations need version and exposure controls beyond the original report.
- Retrieval versus reading: Ground-truth evidence annotations permit retrieval metrics separately from answer correctness. Requiring every annotated evidence session can count retrieval as failed when an updated fact alone answers the question. Conversely, finding all evidence does not ensure successful reasoning. The absence of benchmark-wide shortcut behavior is not proven by the observed association between retrieval and correctness. Recall, answer accuracy, false attribution, and unsupported additions should remain separate outcomes.
- Compression and granularity: The extraction prompts receive only user messages, although the benchmark also asks about assistant-provided information. Consequently, loss under fact/summary replacement reflects that policy and the tested models as well as compression. Round-level advantages vary with reader and token budget; they do not prove a universal optimal memory unit. The useful comparison preserves original values while changing their index keys.
- Numerical inconsistency: Tables 3 and 10 give Stella's round-level
K=VRecall@5/@10 as 0.582/0.692; Table 9 gives 0.660/0.784. Fact expansion is 0.644/0.784 in all three, so Table 9 does not support a consistent improvement in that setting. Session fact-key NDCG also differs: 0.620/0.652 in Tables 3/10 versus 0.520/0.552 in Table 9. Visual inspection confirms these are printed discrepancies. No reconciliation or rerun was performed. Prefer explicitly table-specific estimates; avoid presenting every ablation as a replicated uniform effect. The headline 9.4% recall gain is a relative average, not 9.4 percentage points.
Version boundary, checked 22 September 2026: The pinned repository README announces September 2025 history cleanup to reduce interference with answer correctness and links a cleaned dataset. It also announces LongMemEval-V2 in May 2026. These are later artifacts. Their results and validity are not established by reading the original paper; this note does not claim a full V2 read or quantify the cleanup's effect.
Original reasoning and proposed use: This is a useful engineering scaffold for distinguishing failed acquisition/indexing, failed search, and failed use of available evidence. It does not determine which facts are sufficient to reproduce a person. For an artificial recipient, retain dated source episodes alongside derived facts, preserve who supplied each statement, and test updates without silently treating every inconsistency as replacement. Compare original records, compressed values, and enriched keys under matched retrieval/reader budgets. Use oracle-evidence and no-evidence controls to locate failure, then hold out new probes of already acquired episodes separately from new episodes that test ongoing learning. Keep a development set separate from an untouched, versioned evaluation set; human-review a sample of judge decisions and report answer support separately from permissive benchmark correctness. These are proposed adaptations, not procedures validated by this paper.