ECAI 2025, in Frontiers in Artificial Intelligence and Applications, DOI 10.3233/faia251160, published 21 October 2025.
This comes from the Crossref record (checked 2026-09-26); the arXiv abs page has no journal reference. Read as arXiv
2504.19413 v1 (28 April 2025, the only version), cs.CL with cs.AI, arXiv non-exclusive distribution licence, 23 pages.
The ECAI version was not read or compared. The authors are from Mem0, the company. The author list on the abs page
matches the PDF. For code the paper points to "https://mem0.ai/research", which was not fetched. The
evaluation code and the open-source library are at github.com/mem0ai/mem0 (Apache-2.0), partly read (below).
Provenance: papers/carry_on/chhikara2025_mem0.provenance.json.
What was read
- Paper: all 1,168 lines of the
pdftotext -layouttext of the 23-page PDF. That covers the abstract, sections 1–6, the references, and appendices A (the judge, answer and ChatGPT ingestion prompts), B (Algorithm 1) and C (the baseline descriptions), with Tables 1–2.- Figures 1–4 are images; their captions were read.
- Figures 2 and 3 (the two architectures) were viewed as rendered pages. They show nothing beyond the text.
- Figures 1 and 4 were not viewed. By its caption, Figure 4 plots Table 2's J scores and latencies.
- Evaluation code, commit 393a4fd (29 April 2025): all 16 files of
evaluation/, read in full. This is the first release of the evaluation code, the day after the paper; its code files were unchanged up to 2 May. - Library, commit 1d916c9 (26 April 2025), the last commit before the arXiv submission:
- read in full:
configs/prompts.py,memory/storage.py,graphs/utils.py,memory/graph_memory.py,memory/utils.py; - read in part: lines 1–789 of
memory/main.py, the synchronousMemoryclass. The asynchronous duplicate (lines 790–1565) was only outlined; - not read:
graphs/tools.py. - No library file changed between this commit and the evaluation commit (GitHub compare), so this is also the library at the evaluation release.
- The code was read, not run.
- read in full:
- Data: the original LoCoMo release (snap-research/locomo, commit cbfbc1d). A script counted the questions per category, computed evidence counts and printed sample questions. Mem0's own copy, which their README places on Google Drive, was not fetched.
- Not read: the LoCoMo, A-Mem, Zep, MemGPT and LangMem papers; the result files and the OpenAI memory files that the evaluation code loads (neither is released).
Question
Can a memory layer keep an LLM agent coherent over long multi-session dialogue? The layer extracts salient facts from the conversation, reconciles them with what is stored and retrieves them later. The question is whether it does this more accurately and more cheaply than:
- keeping the whole conversation in context;
- retrieval over raw chunks;
- other memory systems.
And does a graph of entities and relations add to it?
Method
- Mem0, extraction (section 2.1).
- The input is the new message pair, a conversation summary S refreshed in the background, and the last m messages.
- An LLM extracts candidate facts "specifically from the new exchange".
- Mem0, update.
- For each candidate fact, the s most similar stored memories are retrieved.
- "The LLM itself determines which of four distinct operations to execute":
- ADD when nothing equivalent is stored;
- "UPDATE for augmentation of existing memories with complementary information; DELETE for removal of memories contradicted by new information";
- NOOP otherwise.
- Algorithm 1 (appendix B):
- UPDATE replaces the old memory when the new fact carries more information ("Replace with richer information").
- DELETE removes the contradicted memory. The DELETE branch does not add the new fact.
- Settings: m = 10, s = 10, GPT-4o-mini for every LLM step, dense embeddings (model not named).
- Old states. Base Mem0 keeps none: updated and contradicted memories leave the store. The answer prompt adds a read-time rule: "If the memories contain contradictory information, prioritize the most recent memory".
- Mem0g, the graph variant (2.2).
- Structure. A directed labelled graph. Nodes are typed entities with an embedding and a creation timestamp; edges are relation triplets.
- Extraction. Two LLM stages extract entities, then relations. New entities are matched to existing nodes above a similarity threshold t, whose value is not given.
- Conflicts. A conflict detector finds existing relations that may conflict. An LLM "update resolver" decides which become obsolete, "marking them as invalid rather than physically removing them to enable temporal reasoning".
- Retrieval. The subgraph around entities found in the query, plus similarity between the query and each triplet above a threshold.
- Neo4j stores the graph; GPT-4o-mini does the LLM steps.
- Evaluation (section 3).
- Data: LOCOMO. Ten conversations between two people, each of about 600 turns and 26,000 tokens, with about 200 questions each. Four categories: single-hop, multi-hop, temporal and open-domain. The adversarial category was excluded "because ground truth answers were unavailable".
- Measures. F1, BLEU-1, and an LLM judge (J), described as "a separate, more capable LLM".
- The judge prompt (appendix A) tells it to "be generous with your grading - as long as it touches on the same topic as the gold answer, it should be counted as CORRECT".
- J is reported as the mean ± SD of 10 runs.
- Costs. Tokens retrieved per question (counted with cl100k_base), and p50 and p95 latency for the search alone and for the whole answer.
- Baselines, as the paper describes them:
- Earlier systems (LoCoMo, ReadAgent, MemoryBank, MemGPT, A-Mem): "we select the metrics where gpt-4o-mini was used for the evaluation". So their F1 and BLEU-1 come from earlier reports, with no J. A-Mem was re-run at temperature 0 to get J (A-Mem*).
- LangMem: gpt-4o-mini with OpenAI embeddings.
- RAG: chunks of 128 to 8,192 tokens, retrieving the top 1 or 2.
- Full context: the whole conversation in the prompt.
- OpenAI: ChatGPT's memory feature with gpt-4o-mini. Each whole conversation was pasted into a single chat. The resulting memories were used as the full context, "intentionally granting the OpenAI approach privileged access to all memories rather than only question-relevant ones".
- Zep: the commercial memory platform.
Results
Table 1, J (mean of 10 runs; the ± values are 0.11–0.75). The column labels are as printed; see Limits on which questions the columns actually contain.
| Method | "Single Hop" | "Multi-Hop" | "Open Domain" | Temporal |
|---|---|---|---|---|
| A-Mem* | 39.79 | 18.85 | 54.05 | 49.91 |
| LangMem | 62.23 | 47.92 | 71.12 | 23.43 |
| Zep | 61.70 | 41.35 | 76.60 | 49.31 |
| OpenAI | 63.79 | 42.92 | 62.29 | 21.71 |
| Mem0 | 67.13 | 51.15 | 72.93 | 55.51 |
| Mem0g | 65.71 | 47.19 | 75.71 | 58.13 |
- F1 of the earlier systems (no J), in the same column order:
- LoCoMo 25.02 / 12.04 / 40.36 / 18.41;
- ReadAgent 9.15 / 5.31 / 9.67 / 12.60;
- MemoryBank 5.00 / 5.56 / 6.61 / 9.68;
- MemGPT 26.65 / 9.15 / 41.04 / 25.52;
- A-Mem 27.02 / 12.14 / 44.65 / 45.85.
- F1 of Mem0 and Mem0g: Mem0 38.72 / 28.64 / 47.65 / 48.93; Mem0g 38.09 / 24.32 / 49.27 / 51.55.
Table 2, overall J and costs (selected rows):
| Method | Chunk size or memory tokens | Search p50 / p95 (s) | Total p50 / p95 (s) | Overall J |
|---|---|---|---|---|
| Best RAG (k = 2) | 256 | 0.255 / 0.699 | 0.802 / 1.907 | 60.97 |
| Full context | 26031 | – | 9.870 / 17.117 | 72.90 |
| A-Mem | 2520 | 0.668 / 1.485 | 1.410 / 4.374 | 48.38 |
| LangMem | 127 | 17.99 / 59.82 | 18.53 / 60.40 | 58.10 |
| Zep | 3911 | 0.513 / 0.778 | 1.292 / 2.926 | 65.99 |
| OpenAI | 4437 | – | 0.466 / 0.889 | 52.90 |
| Mem0 | 1764 | 0.148 / 0.200 | 0.708 / 1.440 | 66.88 |
| Mem0g | 3616 | 0.476 / 0.657 | 1.091 / 2.590 | 68.44 |
- Full context scores highest: 72.90 overall J, ahead of Mem0g (68.44) and Mem0 (66.88).
- It costs 9.870 s at the median and 17.117 s at p95.
- Per-category scores for full context and RAG are not reported.
- Memory size: Mem0 about 7k tokens per conversation, Mem0g about 14k, Zep over 600k; the raw conversation is about 26k.
- Zep's delay. Zep's retrieval was poor right after ingestion, and "re-running identical searches after a delay of several hours yielded considerably better results". Mem0's graph was ready "in under a minute".
- Headline claims, recomputed from the tables (my computation):
| Claim | Paper | Recomputed |
|---|---|---|
| Mem0 over OpenAI, overall J | 26% | 26.4% |
| Mem0g over Mem0, overall J | around 2% | 2.3% (1.56 points) |
| Mem0 over the best RAG | about 10% | 9.7% |
| Mem0g over the best RAG | around 12% | 12.3% |
| Lower p95 total latency, Mem0 vs full context | 91% (abstract), 92% (section 4.3) | 91.6% |
| Lower p95, Mem0g vs full context | 85% | 84.9% |
| Fewer tokens, Mem0 vs full context | more than 90% | 93.2% |
| Gains in single-hop, temporal, multi-hop (conclusion) | 5%, 11%, 7% | 5.2%, 11.2%, 6.7% |
The last row compares base Mem0 with the best other system in each column. Mem0g's temporal gain over the best other system would be 16.5% (my computation).
Limits
- The evaluated Mem0 is the company's hosted platform, not the open-source library. The evaluation code uses
from mem0 import MemoryClientwith an API key, organisation and project (verified). The platform's code and version are not public, so the Mem0 rows cannot be reproduced from open code (inference). - Extraction was tuned to the benchmark.
- Before ingestion, the code sets instructions for the whole platform project (verified), including:
- "Extract memories only from user messages, not incorporating assistant responses";
- themes such as "Identity and self-acceptance journeys" and "Family planning and parenting".
- These fit LOCOMO's first conversation: Caroline's identity and her adoption plans (my reading of the data). The paper's own ChatGPT prompt example is about "a LGBTQ support group".
- LangMem and OpenAI got generic prompts (inference from the code).
- Before ingestion, the code sets instructions for the whole platform project (verified), including:
- The answer prompts differ (verified in code).
- Mem0, Mem0g, LangMem, Zep and OpenAI get a long prompt with step-by-step rules for timestamps and relative dates.
- RAG and full context get a short one ("Provide the shortest possible answer.").
- The comparison therefore mixes memory with prompt (inference). Full context wins overall anyway.
- The judge is not "more capable". The released judge is
model="gpt-4o-mini"at temperature 0 (verified). That is the same model that extracts the memories and writes the answers. No human check of the judge is reported, and the rubric is generous by design. - Where the ± values come from is unclear. The released
evals.pycalls the judge once per answer; how the 10 runs were made is not in the code. At temperature 0, the SDs (0.11–0.75) can only reflect judge noise, not variation in building the memory or answering (inference). - The category labels look shuffled (my computation; a strong inference).
-
Why Table 2 constrains Table 1. The released
generate_scores.pycomputes the overall score as the mean over all questions. Table 2's overall J should therefore equal the question-weighted mean of Table 1's four columns. -
Counts. The original LoCoMo file has 282, 321, 96 and 841 questions in categories 1–4, and 446 in category 5: 1,540 without category 5 (my count).
-
The fit. Only one assignment reproduces Table 2 for all six systems, to within 0.01:
- "Single Hop" = category 1;
- "Multi-Hop" = category 3;
- "Open Domain" = category 4;
- "Temporal" = category 2.
With single-hop = 4, multi-hop = 1 and open-domain = 3, the weighted means miss Table 2 by 1.80 to 9.67 points.
-
Content, by my reading of the data:
- category 1 is multi-hop: 98% of its questions cite more than one evidence turn (mean 3.13);
- category 4 is single-hop: 5% cite more than one (mean 1.07);
- category 3 is commonsense inference, for example "Would Caroline likely have Dr. Seuss books on her bookshelf?";
- category 2 is temporal.
-
So Table 1's "Single Hop" column holds the multi-hop questions, "Multi-Hop" the 96 commonsense questions, and "Open Domain" the single-hop questions. Temporal is right. The rows copied from other papers may carry the same labels (not checked).
-
Consequences.
- On the real multi-hop questions Mem0 leads: 67.13, against 65.71 for Mem0g and 63.79 for OpenAI.
- On the real single-hop questions Zep leads, with 76.60.
- The paper's reading, that graphs help "open-domain" and do not help "multi-hop", is a reading of mislabelled columns.
-
- Full context wins, and its per-category scores are missing. Whether memory beats full context on temporal or multi-hop questions cannot be seen.
- Abstention is not tested. The excluded adversarial category is 446 of 1,986 questions (22.5%, my count) and the only one that checks abstention. For a memory that deletes and merges, whether it invents answers is exactly the open question (inference).
- Small, synthetic data. Ten conversations, no significance tests, and questions nested within conversations.
- The paper does not say that LOCOMO's conversations were generated by LLM agents.
- Its own appendix C describes LoCoMo's agents with personas and event graphs, which fits my recollection of how the benchmark was built (not verified here).
- Scale claims are not tested. All conversations are about 26k tokens long.
- The claim "As conversation length increases, full-context approaches suffer from exponential growth in computational overhead" rests on RAG chunk sizes, not on conversation length.
- The claim that memory methods "maintain consistent performance regardless of conversation length" was not measured (verified absence).
- Attention cost does not grow exponentially with length (my note).
- How the baselines were run (the code is verified; the consequences are inference):
- Earlier systems. Their F1 and BLEU-1 were copied, with no J. The A-Mem re-run scored 2.92 to 11.31 F1 points below A-Mem's published numbers (my computation from Table 1). The paper does not explain the drop.
- Zep.
- The released ingestion script adds only the first conversation (
if idx == 0:), while the search script queries all ten. - How the reported Zep numbers were produced cannot be reconstructed from the release.
- The authors also report the delayed availability described in Results.
- The released ingestion script adds only the first conversation (
- OpenAI. The memories were copied by hand from the ChatGPT interface ("it processes manually extracted memories from their playground"). The files the code loads are not released.
- LangMem.
- The answer step receives the LangMem agent's reply to the question, not retrieved memories.
- All ingestion calls for a conversation share one agent thread, so the agent's context grows with every message (inference from the code's checkpointer). That may explain its 18-second median latency (guess).
- RAG and full context got the short answer prompt (above).
- The paper and the open-source code differ (verified in the library; the platform may differ):
- Graph conflicts are deleted, not invalidated. The Cypher query runs
DELETE r. - Graph timestamps are ingestion times. Node timestamps come from Neo4j's
timestamp(), not from the conversation dates. - Graph search returns the BM25 top 5, not a list ranked above a threshold.
- Extraction receives only the new messages: no conversation summary and no window of recent messages.
- The update step retrieves 5 similar memories per fact (
limit=5,), not 10, and reconciles all new facts in one call. - History. UPDATE and DELETE are logged to a SQLite history table (old text, new text, event, time). Search never reads it.
- Two prompt lines that matter for a replica:
- The extraction prompt says: "Create the facts based on the user and assistant messages only."
- The library's default answer prompt says: "If no relevant information is found, make sure you don't say no information is found. Instead, accept the question and provide a general response." Where this prompt is used was not checked.
- Graph conflicts are deleted, not invalidated. The Cypher query runs
- Inconsistencies found:
- "consistently outperform". The abstract says the methods "consistently outperform all existing memory systems
across four question categories".
- In Table 1, Zep beats both Mem0 variants in the "Open Domain" column on F1 (49.56) and J (76.60).
- LangMem beats both in the "Multi-Hop" column on BLEU-1 (22.32, against 21.58 and 18.82).
- "scores below 15%". Section 4.1 says "OpenAI notably underperforms, with scores below 15%" on temporal questions. Its temporal J is 21.71; only its F1 (14.04) and BLEU-1 (11.25) are below 15.
- The judge. The paper says "a separate, more capable LLM"; the code uses
gpt-4o-mini. - Graph invalidation in the paper against deletion in the open-source code.
- The category labels (above).
- The latency reduction is 91% in the abstract and conclusion, 92% in section 4.3; recomputed, 91.6%.
- s = 10 in the paper, 5 in the open-source code.
- Checks that passed (my computation):
- every other relative claim, within rounding (table in Results);
- on single-hop, "LangMem and Zep both score around 8% relatively less": 7.3% and 8.1%;
- A-Mem* trails Mem0 by 27.34 J points ("more than 25");
- on open-domain, Zep leads Mem0g by 0.89 and Mem0 by 3.67.
- "consistently outperform". The abstract says the methods "consistently outperform all existing memory systems
across four question categories".
What it means for Kurisutina
- Q3: what a replica with Mem0's design would keep, overwrite and forget, compared with a person (Hirst).
- Keep: short facts, not episodes.
- At encoding, the base LLM decides what is salient and writes it as a sentence.
- Wording, context and most feeling are dropped at that point, unless the extraction prompt asks for them. The LOCOMO instructions did ask for "Emotional states and reactions" (verified).
- Nothing decays with time; the store only grows.
- Overwrite: updates replace in place.
- The prompt's rule is to "keep the fact which has the most information".
- Its example merges "I really like cheese pizza" with "Loves chicken pizza" into "Loves cheese and chicken pizza" (verified).
- A merged memory is a sentence nobody said, with no single source, which is brief 5.3's first field (inference).
- The previous text survives only in the history log, which retrieval does not read.
- Forget: only by contradiction.
- A contradicted memory is deleted. The prompt's example deletes "Loves cheese pizza" on "Dislikes cheese pizza" and does not store the new fact in that step. Algorithm 1 does the same (verified).
- The newest statement wins, at write time and again at read time.
- Compared with Hirst's participants:
- Correction. They repeated wrong personal details at 35 months rather than correcting them, but corrected public facts that the media retold. Mem0 corrects everything a later statement contradicts: personal or public, whoever said it, as long as the extractor attributes it to the user (inference). It is Hirst's "social/cultural reality monitoring" with the most credulous monitor possible.
- Settling. People settled after the first year. Mem0 has no time course: change happens only when contradicting input arrives (inference).
- Emotion. People's reports of feeling were the least consistent. Mem0 keeps a stored feeling unchanged unless it is contradicted (inference).
- The first account. In Hirst, the one-week report was the anchor that made change measurable. Mem0 removes the first version from the store it retrieves from. That runs against brief 6.2 ("Keep first responses separate from later revisions. Never overwrite.") and 5.3 ("an accurate archive must preserve the contradiction").
- The history table is the salvage route. It records every operation (verified), so it could supply the first version and the transitions.
- Against the perfect-retention prediction.
- Mem0 does not retain perfectly: it loses detail at encoding and loses superseded states.
- On LOCOMO, keeping everything (full context, 26k tokens) beat it overall: 72.90 against 66.88 and 68.44.
- At the scale of a few weeks of conversation, extraction is therefore a cost tool, not a fidelity tool (inference).
- Over years of carrying on it may be necessary. Then what is lost must be matched to the person, not to a customer-profile schema (inference).
- Keep: short facts, not episodes.
- Q2: drift.
- The schema. Mem0 has no portrait, but its default extraction prompt defines what a person is: preferences, personal details, plans, activity and service preferences, health and wellness, professional details, and miscellaneous (verified). The base model judges what is salient (inference). A replica using it would remember its new life as a customer profile.
- The benchmark needed more. The evaluation needed benchmark-specific instructions to capture identity, family and mental-health narratives. That suggests the default schema misses much of what LOCOMO's people talk about (inference).
- Other people's statements. Extraction reads user and assistant messages. So an interlocutor's statements can become facts about the self, and someone else's contradicting claim can delete a loaded memory (inference). This makes structural the "memory hacking" risk that Park et al. name and the suggestion effects in the human literature (Lindsay 1990).
- Could Mem0 be a condition in a Q3 test? Yes, as the extract-and-overwrite arm (proposal).
- Setup. The open-source library is the practical route. Whether it runs with a local model was not checked; its factory classes suggest other providers are supported (not verified).
- Procedure. Run it on the replica's new experiences under own, other and empty slots, with the history log on.
- Scoring against the person:
- consistency with the first report at two delays;
- the transitions between delays (repeated, corrected, changed).
- Probes Mem0 was never tested on:
- other people contradicting personal details, set against the person's own tendency to keep or revise them (Hirst: people kept them);
- abstention probes, the category LOCOMO's evaluation excluded.
- Prediction (guess):
- close to the person on facts nobody touches;
- too easily corrected on personal details;
- flat on emotion;
- no forgetting of trivia that the person loses.
- LOCOMO numbers are not evidence about fidelity to an individual. The people are synthetic, the judge is generous and from the same model family, abstention is excluded, and the category labels are probably shuffled.
Cross-references
summaries/carry_on/zhong2023_memorybank.md: decay with rehearsal, the other explicit policy. Its LOCOMO F1 in Table 1 is taken from earlier work.summaries/carry_on/fountas2024_em_llm.md: the perfect-retention end of question 3.summaries/carry_on/park2023_generative_agents.md: memory stream, reflection, "memory hacking".summaries/carry_on/hirst2009_911_memory.md: human consistency and correction over three years.summaries/carry_on/stjacques2013_reactivation.md: human memories updated when reactivated alongside new material.summaries/carry_on/venkit2026_companion_drift.md: three memory settings at chance on user-state changes.summaries/memory/wu2025_longmemeval.md: knowledge updates; keeping the original conversation as the value while adding extracted facts as retrieval keys helped. Mem0 does the opposite and replaces the values with facts.summaries/memory/lindsay1990_source_suggestions.md,summaries/memory/johnson1993_source_monitoring.md: suggestion and source confusion in people.- Brief v1.5, sections 5.3 (provenance fields and ignorance states) and 6.2 (design rules).