arXiv 2502.12110, version 11 (8 October 2025); eleven versions since 17 February 2025. NeurIPS 2025 per the arXiv
comments. Rutgers University and the AIOS Foundation. arXiv non-exclusive distribution licence. Code: two GitHub
repositories (not read). Provenance: papers/carry_on/xu2025_amem.provenance.json.
What was read
All 1,673 lines of pdftotext -layout output (28 pages): text, Tables 1–8, captions and extracted labels of Figures
1–5, Appendices A–B (baselines, metrics, extra results, k settings, the three prompts, a Q/A example) and the NeurIPS
checklist. Figures are images; the numbers in Figure 3 were available as extracted labels. The code was not read.
Question
Can an LLM agent's memory organise itself? Each new memory becomes a note that links to related notes and evolves their descriptions. Does this answer long-conversation questions better, and more cheaply, than full-context and earlier memory systems?
Method
- Note construction. Each interaction becomes a note $m_i = {c_i, t_i, K_i, G_i, X_i, e_i, L_i}$:
- verbatim content $c_i$ and timestamp $t_i$;
- LLM-generated keywords $K_i$, tags $G_i$ and a one-sentence context $X_i$;
- an embedding $e_i$ (all-MiniLM-L6-v2);
- links $L_i$.
- Link generation. Take the top-k notes by cosine similarity; the LLM decides which to link.
- Memory evolution. For each neighbour, the LLM may rewrite its context and tags in the light of the new note. "The evolved memory $m_j^*$ then replaces the original memory $m_j$." The content $c_j$ is not in the list of rewritable fields (prompt B.3).
- Retrieval. Top-k notes by cosine similarity to the query; linked notes come along.
- Evaluation.
- Datasets: LoCoMo (5 question categories, including adversarial) and DialSim (TV-show dialogues).
- Six backbones: GPT-4o-mini, GPT-4o, Qwen2.5 1.5B/3B and Llama 3.2 1B/3B. Three more in the appendix.
- Baselines: full context ("LoCoMo"), ReadAgent, MemoryBank and MemGPT.
- Metrics: F1, BLEU-1, ROUGE, METEOR and SBERT. No error bars (checklist: API cost).
- k was set per category and model (Table 8: 40–50 for GPT, mostly 10 otherwise).
Results
-
LoCoMo with GPT-4o-mini (F1, Table 1):
System Multi Hop Temporal Open Domain Single Hop Adversarial A-Mem 27.02 45.85 12.14 44.65 50.03 Full context 25.02 18.41 12.04 40.36 69.23 MemGPT 26.65 25.52 9.15 41.04 43.29 - A-Mem gains most on Temporal.
- Full context wins on Adversarial (unanswerable) questions: 69.23 against 50.03 for A-Mem.
- With GPT-4o, full context also wins Single Hop (61.56 against 48.43).
- On the small open models A-Mem leads every category.
-
Cost. About 1,200–2,500 tokens per answer against 16,900 for full context, an 85–93% reduction.
-
DialSim. F1 3.45 against 2.55 for full context and 1.18 for MemGPT.
-
Ablation (Table 3, GPT-4o-mini F1). Adding link generation, then evolution:
Configuration Multi Hop Temporal Neither 9.65 24.55 Links only 21.35 31.24 Full A-Mem 27.02 45.85 -
Other results. Performance levels off beyond k ≈ 30–40. t-SNE plots show tighter clusters with evolution.
Limits
- The evaluation is question answering about facts in dialogues. Nothing checks whether an evolved description is faithful to what happened, or how descriptions change over many rewrites.
- k is chosen per category and model on the evaluation data (inferred from Table 8 and "adjusting this parameter for specific categories to optimize performance"). There are no error bars and no repeated runs.
- Evolution overwrites the old context and tags, and no version history is described.
- The prompt text offers "strengthen, update_neighbor", but the JSON example lists "strengthen", "merge", "prune". The last two are never described.
- Inconsistencies found:
- The text calls the Temporal column "Multi-Hop".
- The main text claims "superior performance in Multi-Hop tasks achieves at least two times better performance". Multi Hop in Table 1 is 27.02 against 26.65 (MemGPT) and 25.02; only Temporal (45.85 against 25.52/18.41) doubles.
- Appendix A.3 cites "ROUGE-L … 44.27 in Multi-Hop … LoComo's 18.09", "METEOR … 23.43 versus 7.61" and "SBERT … 70.49 versus 52.30". All of these are Table 5/6 Temporal values.
- Likewise "Qwen2.5-15b … ROUGE-L … 27.23" is Qwen2.5-1.5B Temporal.
- Duplicated rows:
- The ReadAgent row in Table 1 is identical for Qwen2.5-3B and Llama 3.2-3B, differing only in token length.
- Table 5's Llama 3.2-3B ReadAgent row repeats Table 1 F1/BLEU values (2.47, 1.78, 3.01 …) in the ROUGE columns.
- "Primarily employ k=10", but GPT runs used k = 40–50 in every category (Table 8).
- Table 4 retrieval times of 0.31–3.70 µs for 1,000–1,000,000 stored vectors are implausibly fast for exact cosine search over a million 384-dimensional embeddings (a guess from arithmetic: about 400 million multiply-adds per query). The method is not described.
- The text calls the Temporal column "Multi-Hop".
- Checked:
- DialSim +35% (3.45/2.55) and +192% (3.45/1.18).
- The 85–93% token reduction (2,520 and 1,200 against about 16,900).
Relation to Mem0 (resolves a subagent 4 finding; verified here)
-
Mem0's Table 1 rows for full context, ReadAgent, MemoryBank, MemGPT and A-Mem are A-Mem's GPT-4o-mini numbers with the column labels permuted:
Mem0 label A-Mem label "Single Hop" Multi Hop "Multi-Hop" Open Domain "Open Domain" Single Hop "Temporal" Temporal All five rows match exactly under that mapping (script comparison of both tables).
-
Per subagent 4's count, A-Mem's column labels match LoCoMo's categories by content:
- category 1 is multi-hop;
- category 3 is commonsense or open domain;
- category 4 is single-hop.
-
So the error is Mem0's relabelling, not LoCoMo's.
-
A-Mem's own text has a separate mislabel (Temporal described as Multi-Hop, above).
What it means for Kurisutina
- Q3: LLM-driven reconsolidation, without the person.
- A-Mem rewrites how old memories are described whenever a related new memory arrives.
- This is structurally like human updating: the content is kept, and its framing and associations change (compare Speer: 73–75% of original details kept while feeling and framing shift).
- It helped retrieval, and temporal questions most (31.24 to 45.85 F1). Re-contextualising old memories has real functional value.
- But the base model writes the new framing (inferred). Evolution uses a generic prompt: "update the context and
tags … based on the understanding of these memories". Every rewrite imports the base model's reading of the
person's past.
- Over many rewrites, a replica's account of its own history would drift toward how the model frames things. That is exactly the change the goal rules out.
- Nothing in the paper measures it.
- Architecture pattern worth keeping (proposal):
- an immutable content layer (verbatim episode, timestamp, encoding-time affect; see the FAB summaries);
- a mutable interpretive layer (context, links, current affect), versioned, never overwritten;
- rewrites of the interpretive layer driven by the person's own appraisal habits, or at least audited against them; not by a generic model prompt.
- Q3 test (proposal).
- Feed a replica with an A-Mem-style store a long sequence of experiences.
- Track, for a fixed set of early memories, the drift of their evolved context:
- embedding distance from the original description;
- judge-free style markers (Breithaupt);
- affect.
- Compare with the person's own re-descriptions of the same events at the same intervals.
- Adversarial questions matter for a replica. Full context beat A-Mem on unanswerable questions (69.23 against 50.03). A memory system that retrieves plausible neighbours is more prone to answering what never happened. For a replica, confabulated personal history is a Q3 failure. The adversarial-question probe should stay in any memory evaluation; Mem0 dropped it.
Cross-references
summaries/carry_on/chhikara2025_mem0.md: overwrite-on-update memory; the relabelled table.summaries/carry_on/zhong2023_memorybank.md: forgetting-curve memory.summaries/carry_on/fountas2024_em_llm.md: event-segmented memory.summaries/carry_on/speer2021_positive_meaning_memory.md: human updating at recall.summaries/carry_on/breithaupt2024_retelling_novelty.md: judge-free style markers.