Kurisutina

Benchmarking chat assistants on long-term interactive memory

arXiv 2410.10813, version 2 (4 March 2025), published at ICLR 2025. UCLA, Tencent AI Lab and UC San Diego. CC BY 4.0. Benchmark and code: github.com/xiaowu0162/LongMemEval (not read). Provenance: papers/carry_on/wu2024_longmemeval.provenance.json.

What was read

All 1,821 lines of pdftotext -layout output (28 pages): main text, Tables 1–11, the figures' extracted labels and captions, and Appendices A–E (construction, statistics, judge prompts and meta-evaluation, the human study of commercial systems, the unified view, implementation prompts, extended analyses). Figures were not rendered. The dataset and code were not read.

Question

How well do chat assistants remember what a user told them over long, realistic histories? Which memory designs (what to store, how to index, how to query, how to read) help?

Method

  • Benchmark. 500 hand-curated questions of seven types:

    • single-session-user, single-session-assistant and single-session-preference;
    • multi-session;
    • knowledge-update;
    • temporal-reasoning;
    • abstention: 30 false-premise questions ("How many fish are in my 30-gallon tank?" when no such tank was mentioned).

    Together they cover five abilities: extraction, multi-session reasoning, knowledge updates, temporal reasoning, abstention.

  • Construction.

    • A 164-attribute user ontology feeds Llama 3 70B, which writes invented user backgrounds.
    • Question proposals were human-filtered to about a 5% yield.
    • Evidence statements are embedded indirectly in LLM self-chat sessions; about 70% of sessions were human-edited.
    • Sessions are mixed into haystacks of unrelated sessions (ShareGPT, UltraChat and simulated) with timestamps:
      • LongMemEval-S: about 115k tokens, around 50 sessions;
      • LongMemEval-M: 500 sessions, about 1.5M tokens.
  • Judge. GPT-4o with per-type prompts; 0.90–1.00 agreement with experts on 30 questions per type. The knowledge-update prompt accepts answers that also mention the old value.

  • Framework. Indexing, then retrieval, then reading, with four control points: value granularity, key design, query, and reading strategy.

Results

  • Commercial memory assistants (97 questions, short 3–6-session histories, human-operated in August 2024): with GPT-4o, ChatGPT reached 0.577 and Coze 0.330, against 0.918 when GPT-4o reads the whole history offline.
    • ChatGPT "often modify this information when it compresses the history, resulting in information loss".
    • Coze often failed to record information given in passing.
  • Long context is not memory. At about 115k tokens, accuracy fell 30–66% against reading only the evidence sessions:
    • GPT-4o from 0.870 to 0.606;
    • Llama 3.1 70B from 0.744 to 0.334.
  • Design findings:
    • Storing rounds beats storing sessions.
    • Replacing raw content with summaries or extracted facts loses information and lowers accuracy, except for multi-session aggregation.
    • Keeping the raw value but expanding the key with extracted user facts raises recall by 9.4% and accuracy by 5.4%.
    • Time-aware query expansion raises temporal recall by 6.8–11.3%, but only with a strong model extracting the range. Llama 8B invents time ranges.
    • Extract-then-reason reading (Chain-of-Note) plus JSON formatting is worth up to 10 points even with perfect retrieval.
  • Errors. Even with the best design, 43–54% of errors (my computation from Figure 14) had correct retrieval and wrong reading. The authors say 40–50%.
    • Most answers that were right despite "failed" retrieval are knowledge updates where only the new value was retrieved.

Limits

  • Users, backgrounds and conversations are LLM-generated, then human-edited. There are no real people and no real life histories.
  • The memory tested is an assistant's memory of a user, of facts rather than experiences.
  • The judge is a single closed model (GPT-4o), and the commercial study is small and manual.
  • Inconsistencies found:
    • Same configuration, different numbers. Stella V5 with K = V gives Recall@5 0.582 (rounds) and 0.706 (sessions) in Table 3, but 0.660 and 0.720 in Table 9. The K = V + fact rows agree (0.644 and 0.732).
    • Appendix E.4 lists "example 1, 2, 4" as questions without temporal reference, but example 2 is the one with a clear range. The intended list is probably 1, 3, 4.
    • The share of errors with correct retrieval computes to 43–54% from Figure 14; the text gives 40–50%.
  • Checked:
    • The long-context drops in Figure 3b (e.g. 1 − 0.606/0.870 = 30.3%).
    • The commercial drops of 37% and 64%.
    • Each Figure 14 pie sums to 100%.

What it means for Kurisutina

  • Q3 test shape. The five abilities are a ready checklist for the replica's memory of its own life:

    • recall a detail;
    • aggregate across episodes;
    • know what changed and when;
    • reason about time;
    • refuse to recall what never happened.

    The abstention type (false premise) is the confabulation probe that Mem0 dropped and A-Mem handled worse than full context. It should be standard in every Q3 evaluation (proposal).

  • Knowledge update, done the human way. The benchmark's judge accepts an answer that mentions the old value as long as the new one is given.

    • For a replica that is the right target: a person knows their family went to Hawaii last month and Paris last week.
    • Systems that retrieve only the update (the error analysis) or overwrite on compression (ChatGPT) lose the history a person keeps.
  • Store raw episodes; index with extracted facts; never replace. Summaries and fact extraction as values lost information. As keys over raw values they helped.

    • This matches the architecture proposed from A-Mem: an immutable content layer plus a derived, versioned interpretive layer. The derived layer should be allowed to guide retrieval, never to replace what happened.
  • Commercial memory already shows drift by compression. ChatGPT's memory altered stored details as it condensed history.

    • This is the LLM analogue of Acerbi's condensation bias, and the mechanism by which a replica's past could quietly change. It needs the same versioning and audit.
  • Time needs structure. Temporal questions need timestamps as data, not text. A small local model invents date ranges, so the replica's own timeline should be explicit metadata (proposal).

Cross-references

  • summaries/carry_on/chhikara2025_mem0.md, summaries/carry_on/xu2025_amem.md, summaries/carry_on/zhong2023_memorybank.md: memory systems evaluated here or on LoCoMo.
  • summaries/carry_on/acerbi2023_llm_transmission.md: what condensation keeps.
  • summaries/carry_on/diamond2020_real_world_recall.md: human recall accuracy and quantity.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.