Kurisutina

A-Mem: agentic memory for LLM agents

arXiv 2502.12110, version 11 (8 October 2025); eleven versions since 17 February 2025. NeurIPS 2025 per the arXiv comments. Rutgers University and the AIOS Foundation. arXiv non-exclusive distribution licence. Code: two GitHub repositories (not read). Provenance: papers/carry_on/xu2025_amem.provenance.json.

What was read

All 1,673 lines of pdftotext -layout output (28 pages): text, Tables 1–8, captions and extracted labels of Figures 1–5, Appendices A–B (baselines, metrics, extra results, k settings, the three prompts, a Q/A example) and the NeurIPS checklist. Figures are images; the numbers in Figure 3 were available as extracted labels. The code was not read.

Question

Can an LLM agent's memory organise itself? Each new memory becomes a note that links to related notes and evolves their descriptions. Does this answer long-conversation questions better, and more cheaply, than full-context and earlier memory systems?

Method

  • Note construction. Each interaction becomes a note $m_i = {c_i, t_i, K_i, G_i, X_i, e_i, L_i}$:
    • verbatim content $c_i$ and timestamp $t_i$;
    • LLM-generated keywords $K_i$, tags $G_i$ and a one-sentence context $X_i$;
    • an embedding $e_i$ (all-MiniLM-L6-v2);
    • links $L_i$.
  • Link generation. Take the top-k notes by cosine similarity; the LLM decides which to link.
  • Memory evolution. For each neighbour, the LLM may rewrite its context and tags in the light of the new note. "The evolved memory $m_j^*$ then replaces the original memory $m_j$." The content $c_j$ is not in the list of rewritable fields (prompt B.3).
  • Retrieval. Top-k notes by cosine similarity to the query; linked notes come along.
  • Evaluation.
    • Datasets: LoCoMo (5 question categories, including adversarial) and DialSim (TV-show dialogues).
    • Six backbones: GPT-4o-mini, GPT-4o, Qwen2.5 1.5B/3B and Llama 3.2 1B/3B. Three more in the appendix.
    • Baselines: full context ("LoCoMo"), ReadAgent, MemoryBank and MemGPT.
    • Metrics: F1, BLEU-1, ROUGE, METEOR and SBERT. No error bars (checklist: API cost).
    • k was set per category and model (Table 8: 40–50 for GPT, mostly 10 otherwise).

Results

  • LoCoMo with GPT-4o-mini (F1, Table 1):

    System Multi Hop Temporal Open Domain Single Hop Adversarial
    A-Mem 27.02 45.85 12.14 44.65 50.03
    Full context 25.02 18.41 12.04 40.36 69.23
    MemGPT 26.65 25.52 9.15 41.04 43.29
    • A-Mem gains most on Temporal.
    • Full context wins on Adversarial (unanswerable) questions: 69.23 against 50.03 for A-Mem.
    • With GPT-4o, full context also wins Single Hop (61.56 against 48.43).
    • On the small open models A-Mem leads every category.
  • Cost. About 1,200–2,500 tokens per answer against 16,900 for full context, an 85–93% reduction.

  • DialSim. F1 3.45 against 2.55 for full context and 1.18 for MemGPT.

  • Ablation (Table 3, GPT-4o-mini F1). Adding link generation, then evolution:

    Configuration Multi Hop Temporal
    Neither 9.65 24.55
    Links only 21.35 31.24
    Full A-Mem 27.02 45.85
  • Other results. Performance levels off beyond k ≈ 30–40. t-SNE plots show tighter clusters with evolution.

Limits

  • The evaluation is question answering about facts in dialogues. Nothing checks whether an evolved description is faithful to what happened, or how descriptions change over many rewrites.
  • k is chosen per category and model on the evaluation data (inferred from Table 8 and "adjusting this parameter for specific categories to optimize performance"). There are no error bars and no repeated runs.
  • Evolution overwrites the old context and tags, and no version history is described.
  • The prompt text offers "strengthen, update_neighbor", but the JSON example lists "strengthen", "merge", "prune". The last two are never described.
  • Inconsistencies found:
    • The text calls the Temporal column "Multi-Hop".
      • The main text claims "superior performance in Multi-Hop tasks achieves at least two times better performance". Multi Hop in Table 1 is 27.02 against 26.65 (MemGPT) and 25.02; only Temporal (45.85 against 25.52/18.41) doubles.
      • Appendix A.3 cites "ROUGE-L … 44.27 in Multi-Hop … LoComo's 18.09", "METEOR … 23.43 versus 7.61" and "SBERT … 70.49 versus 52.30". All of these are Table 5/6 Temporal values.
      • Likewise "Qwen2.5-15b … ROUGE-L … 27.23" is Qwen2.5-1.5B Temporal.
    • Duplicated rows:
      • The ReadAgent row in Table 1 is identical for Qwen2.5-3B and Llama 3.2-3B, differing only in token length.
      • Table 5's Llama 3.2-3B ReadAgent row repeats Table 1 F1/BLEU values (2.47, 1.78, 3.01 …) in the ROUGE columns.
    • "Primarily employ k=10", but GPT runs used k = 40–50 in every category (Table 8).
    • Table 4 retrieval times of 0.31–3.70 µs for 1,000–1,000,000 stored vectors are implausibly fast for exact cosine search over a million 384-dimensional embeddings (a guess from arithmetic: about 400 million multiply-adds per query). The method is not described.
  • Checked:
    • DialSim +35% (3.45/2.55) and +192% (3.45/1.18).
    • The 85–93% token reduction (2,520 and 1,200 against about 16,900).

Relation to Mem0 (resolves a subagent 4 finding; verified here)

  • Mem0's Table 1 rows for full context, ReadAgent, MemoryBank, MemGPT and A-Mem are A-Mem's GPT-4o-mini numbers with the column labels permuted:

    Mem0 label A-Mem label
    "Single Hop" Multi Hop
    "Multi-Hop" Open Domain
    "Open Domain" Single Hop
    "Temporal" Temporal

    All five rows match exactly under that mapping (script comparison of both tables).

  • Per subagent 4's count, A-Mem's column labels match LoCoMo's categories by content:

    • category 1 is multi-hop;
    • category 3 is commonsense or open domain;
    • category 4 is single-hop.
  • So the error is Mem0's relabelling, not LoCoMo's.

  • A-Mem's own text has a separate mislabel (Temporal described as Multi-Hop, above).

What it means for Kurisutina

  • Q3: LLM-driven reconsolidation, without the person.
    • A-Mem rewrites how old memories are described whenever a related new memory arrives.
    • This is structurally like human updating: the content is kept, and its framing and associations change (compare Speer: 73–75% of original details kept while feeling and framing shift).
    • It helped retrieval, and temporal questions most (31.24 to 45.85 F1). Re-contextualising old memories has real functional value.
  • But the base model writes the new framing (inferred). Evolution uses a generic prompt: "update the context and tags … based on the understanding of these memories". Every rewrite imports the base model's reading of the person's past.
    • Over many rewrites, a replica's account of its own history would drift toward how the model frames things. That is exactly the change the goal rules out.
    • Nothing in the paper measures it.
  • Architecture pattern worth keeping (proposal):
    • an immutable content layer (verbatim episode, timestamp, encoding-time affect; see the FAB summaries);
    • a mutable interpretive layer (context, links, current affect), versioned, never overwritten;
    • rewrites of the interpretive layer driven by the person's own appraisal habits, or at least audited against them; not by a generic model prompt.
  • Q3 test (proposal).
    • Feed a replica with an A-Mem-style store a long sequence of experiences.
    • Track, for a fixed set of early memories, the drift of their evolved context:
      • embedding distance from the original description;
      • judge-free style markers (Breithaupt);
      • affect.
    • Compare with the person's own re-descriptions of the same events at the same intervals.
  • Adversarial questions matter for a replica. Full context beat A-Mem on unanswerable questions (69.23 against 50.03). A memory system that retrieves plausible neighbours is more prone to answering what never happened. For a replica, confabulated personal history is a Q3 failure. The adversarial-question probe should stay in any memory evaluation; Mem0 dropped it.

Cross-references

  • summaries/carry_on/chhikara2025_mem0.md: overwrite-on-update memory; the relabelled table.
  • summaries/carry_on/zhong2023_memorybank.md: forgetting-curve memory.
  • summaries/carry_on/fountas2024_em_llm.md: event-segmented memory.
  • summaries/carry_on/speer2021_positive_meaning_memory.md: human updating at recall.
  • summaries/carry_on/breithaupt2024_retelling_novelty.md: judge-free style markers.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.