AAAI 2024: Proceedings of the AAAI Conference on Artificial Intelligence 38(17): 19724–19731, DOI
10.1609/aaai.v38i17.29946. This comes from the Crossref record (checked 2026-09-26), and the Mem0 paper cites the same
volume and pages. The arXiv abs page has no journal reference. Read as arXiv 2305.10250 v3 (21 May 2023; v1 of 17 May
and v2 of 18 May 2023 not read), cs.CL with cs.AI, arXiv non-exclusive distribution licence. The abs page says "10 pages";
the PDF has 11, the last being references. The AAAI version was not read or compared with v3. Sun Yat-Sen University,
Harbin Institute of Technology and KTH. The author list on the abs page matches the PDF. Code and data are at
github.com/zhongwanjun/MemoryBank-SiliconFriend (MIT licence), partly read (below).
Provenance: papers/carry_on/zhong2023_memorybank.provenance.json.
What was read
- Paper: all 678 lines of the
pdftotext -layouttext of the 11-page PDF. That covers the abstract, sections 1–6, the footnotes, Tables 1–2 and the references.- Figures 1–4 are vector drawings. Their text (Figure 1's labels, the dialogues in Figures 2–4) was extracted and read; the drawings were not viewed.
- In Figure 2 the two columns interleave in the text, so which bubble belongs to which model is partly ambiguous.
- Code, not part of the paper, at the repository's last commit, cf61c41 (24 May 2023):
- read in full:
forget_memory.py(the forgetting rule),local_doc_qa.py(retrieval without forgetting),summarize_memory.py(the summary and portrait prompts),build_memory_index.py,model_config.py,utils/memory_utils.py,utils/prompt_utils.py,utils/sys_args.py, the three demo scripts, the four launch scripts, the README and the English probing questions; - counted by script, not read line by line: the English and Chinese simulated memory banks and the Chinese probing questions;
- not read: the training scripts,
model_utils.py, the app modules, the Chinese README and the LoRA checkpoints. - The code was read, not run.
- read in full:
- Not available: the real users' conversations, the 38k psychological dialogues, the evaluation driver and the annotators' labels. None of them is in the repository.
Question
Can an LLM chatbot get a long-term memory that recalls past conversations, builds a picture of the user's personality, and forgets and strengthens memories in a human-like way? The setting is an AI companion ("SiliconFriend") on a closed model (ChatGPT) and two open ones (ChatGLM, BELLE), in English and Chinese.
Method
- Storage, three layers (section 2.1):
- every turn, with a timestamp;
- an event summary per day, then a global summary of the daily ones. Prompt: "Summarize the events and key information in the content [dialog/events]";
- a personality portrait per day, then a global portrait. Daily prompt: "Based on the following dialogue, please summarize the user’s personality traits and emotions.[dialog]". Global prompt: "Please provide a highly concise and general summary of the user’s personality".
- Retrieval (2.2):
- a dual-encoder dense retriever as in DPR, indexed with FAISS;
- each turn and each event summary is a memory piece, and the current conversation context is the query;
- in practice LangChain, with MiniLM for English and Text2vec for Chinese;
- the reply prompt receives the retrieved memories, the global portrait and the global event summary;
- how many memories are retrieved is not stated.
- Forgetting and strengthening (2.3):
- Retention is R = exp(−t/S), where t is the time since learning and S the memory strength.
- S is discrete. It is set to 1 "upon its first mention in a conversation".
- "When a memory item is recalled during conversations, it will persist longer in memory. We increase S by 1 and reset t to 0, hence forget it with a lower probability."
- Not stated in the paper:
- the unit of t;
- how R is used: as a probability of deletion, a threshold or a retrieval weight;
- what counts as a recall;
- which layers can be forgotten;
- whether "forgotten" means deleted.
- The authors call it "an exploratory and highly simplified memory updating model", and add: "The forgetting curve will look different for different people and different types of information."
- The motivation given: "Forgetting less important memory pieces that are long time ago and have not been recalled much can make the AI companion more natural."
- SiliconFriend (section 3):
- ChatGPT is used untuned.
- ChatGLM (6.2B) and BELLE (7B, based on LLaMA) are LoRA-tuned (rank 16, 3 epochs, one A100) on 38k psychological dialogues, "parsed from online sources".
- Evaluation (section 4):
- Qualitative, real users: "we have developed an online platform for SiliconFriend and collected real-time conversations from actual users." The number of users is not given. Three examples are shown (Figures 2–4).
- Quantitative, simulated users:
- 15 virtual users; ChatGPT generated their names, personalities and interests;
- 10 days of conversation, with at least two topics per day. "Conversations are synthesized by users acted by ChatGPT based on predefined topics and user personalities." Who wrote the AI turns is not stated;
- "we manually write 194 probing questions (97 in English and 97 in Chinese)";
- "human annotators" label each answer. Their number, their agreement and any blinding are not reported.
- Measures:
- retrieval accuracy (0 or 1);
- correctness (0, 0.5 or 1);
- coherence (0, 0.5 or 1);
- a ranking score s = 1/r across the three variants.
- Conditions. The only conditions are the three SiliconFriend variants. There is no condition without memory, none without forgetting, and no other memory system.
Results
Table 2 (simulated users):
| Language | Variant | Retrieval acc. | Correctness | Coherence | Ranking |
|---|---|---|---|---|---|
| English | ChatGLM | 0.809 | 0.438 | 0.68 | 0.498 |
| English | BELLE | 0.814 | 0.479 | 0.582 | 0.517 |
| English | ChatGPT | 0.763 | 0.716 | 0.912 | 0.818 |
| Chinese | ChatGLM | 0.84 | 0.418 | 0.428 | 0.51 |
| Chinese | BELLE | 0.856 | 0.603 | 0.562 | 0.565 |
| Chinese | ChatGPT | 0.711 | 0.655 | 0.675 | 0.758 |
-
The authors' reading:
- the ChatGPT variant is best overall;
- the open-source variants retrieve as well or better but answer worse. The authors write: "This might be attributed to the inferior overall abilities of the base models";
- ChatGLM and ChatGPT are better in English, BELLE in Chinese.
-
Qualitative examples:
- Figure 2: after a break-up, the tuned ChatGLM gives more emotional support than the base ChatGLM.
- Figure 3: BELLE recalls a book it recommended and a quicksort program the user asked for. It also correctly denies that the two wrote a heap sort together.
- Figure 4: ChatGPT tailors weekend suggestions to two users' portraits.
-
Nothing on forgetting. No experiment reports:
- forgetting switched on against off;
- a forgetting rate;
- what was lost.
The introduction's third contribution, "Applicability with and without memory forgetting mechanism", has no result behind it.
Limits
- The forgetting rule is never tested (see Results). The portrait is not tested either: nothing measures whether it describes the user correctly.
- The released code does not implement the paper's rule (verified by reading the code at commit cf61c41; not run).
-
Strength speeds up forgetting. The formula is
return math.exp(-t / 5*S). Python reads this as exp(−t·S/5), so greater strength means faster forgetting. The docstring says the opposite: "The higher the memory strength, the slower the rate of forgetting,". Values, with t in days (my computation):t (days) S Paper, exp(−t/S) Code, exp(−t·S/5) 1 1 .368 .819 1 2 .607 .670 1 5 .819 .368 7 1 .001 .247 30 1 .000 .002 Under the code, a memory retrieved four times (S = 5) keeps .368 a day after its last retrieval. One never retrieved keeps .819 a day after it was made.
-
The clock. t counts days since the last retrieval, or since the turn's own date if it was never retrieved.
-
Deletion is random and permanent. Each turn gets a random draw:
if random.random() > retention_probability:. A turn that fails the draw is removed and the memory file is rewritten. -
The draw repeats. It runs whenever the memory index is rebuilt. The demos do that each time a user enters their name, and the draw covers every user in the file. Survival therefore depends on how often anyone logs in, not only on time. For a turn never retrieved, five daily draws leave exp(−3) = .050 by day 5; one draw on day 5 leaves .368 (my computation).
-
Not everything can be forgotten.
- Day summaries are never drawn. A day's summary is removed only when all of that day's turns are gone.
- The daily and global portraits are never touched by the forgetting code.
-
What counts as a recall. A turn is "recalled" when it is among the top k retrieved for a user message (k = 2 by default), whether or not the answer uses it. Summaries can never be strengthened, because their ids never match the lookup.
-
Forgetting is off by default:
enable_forget_mechanism: bool = field(default=False).- No launch script turns it on. The BELLE script passes
--enable_forget_mechanism Falseand calls a script that is not in the repository. - With the flag on, the command-line demo imports
forget_memory_new, which is not in the repository either.
- No launch script turns it on. The BELLE script passes
-
The ChatGPT variant has no forgetting path at all. It builds a llama_index vector index of whole days. The prompt then receives the
.responseof the index's query, which in llama_index of that period is a generated answer, not retrieved turns (the last point is inference). -
Inference: the Table 2 systems probably ran without forgetting. At least the ChatGPT variant, the best one, could not have used it with the released code. The paper's own rule points the same way. With t in days, a turn never recalled would keep e^−7 ≈ .001 over the week between Table 1's memory (3 May) and its probe (10 May), which fits poorly with successful week-old recall (my computation).
-
- The simulated test measures retrieval over a fixed store, not memory formation. One store per language was built before testing and shared by the three variants. Its AI turns were therefore not produced by the system under test (inference). ChatGPT generated both the users and their personalities.
- Retrieval accuracy differs between variants, although the paper describes one retriever.
- The code explains part of this: the ChatGPT variant uses a different retrieval stack (llama_index). Its embedding model was not checked.
- ChatGLM and BELLE share one retriever in the code and still differ (0.809 against 0.814 in English, 0.84 against 0.856 in Chinese). Either the query carries dialogue context that differs between models, or the labels are noisy (guess).
- Judges and users are not described.
- There is no count of annotators or of real users, and no agreement statistic.
- Consent and ethics review are not mentioned. This was checked by searching the full text.
- Tuning is confounded with the base model. Only the open-source variants were tuned. The evidence for empathy is one example against the untuned base model.
- An evaluation prompt in the code may discourage admitting a gap. A meta-prompt in
prompt_utils.pyends: "Please answer my question according to the memory and it's forbidden to say sorry." That works against brief 5.3's first-class output, "I don’t remember that" (inference). The evaluation driver is not released, so whether this prompt was used is unknown. - Code observation, not run. In the forgetting loader the per-user filter is commented out. Each user's index is
therefore built from all users' memories, so other users' conversations could be retrieved into a user's prompt
(inference).
local_doc_qa.py, the loader without forgetting, filters by user. - Inconsistencies found:
- Number of probing questions. The paper has 194 (97 per language). The README says "we manually craft 100
probing questions". The released files hold 100 per language (my count).
- Table 2 fits 97 questions but not 100 (my computation). With one label per question and 97 questions, every cell is attainable except English ChatGLM retrieval, 0.809: 78/97 = .804 and 79/97 = .814.
- With 100 questions, 14 of the 18 cells in the first three columns are not attainable.
- So the released questions are probably not the scored set (inference).
- The formula. The paper's rule and the code's rule differ (above).
- "Qualitative" and "quantitative". Section 4 announces "a qualitative analysis that uses simulated long-term dialog history and 194 memory probing questions". Section 4.2 then presents it as the quantitative analysis.
- "relative significance". The abstract says the AI can "forget and reinforce memory based on time elapsed and the relative significance of the memory". The rule has no significance term; the only candidate is the recall count (my reading).
- "Spacing Effect". The bullet under that name describes savings instead: "Ebbinghaus discovered that relearning information is easier than learning it for the first time." The rule does not depend on the spacing of recalls either, since S rises by 1 whatever the interval (my note).
- The language claims hold only in part.
- ChatGLM is better in English on correctness and coherence, but not on retrieval (0.809 against 0.84) or ranking (0.498 against 0.51).
- BELLE is better in Chinese except on coherence (0.582 in English against 0.562).
- The portrait prompts differ from the code.
- The code's daily prompt adds "devise response strategies based on your speculation".
- The code's global prompt reads "Please provide a highly concise and general summary of the user's personality and the most appropriate response strategy for the AI lover, summarized as:".
- Checks that passed (my computation):
- The ranking scores sum to 1.833 in both languages, as strict rankings of 1, 1/2 and 1/3 require.
- The released English store matches the described scale: 15 users, 10 days (27 April to 6 May 2023), 566 turns (17 to 52 per user). Table 1's example (Gary, stress, 3 May) is in it.
- The paper reports no averages or headline percentages, so there were none to recompute.
- Number of probing questions. The paper has 194 (97 per language). The README says "we manually craft 100
probing questions". The released files hold 100 per language (my count).
What it means for Kurisutina
- Q3: what a replica with MemoryBank's design would keep, overwrite and forget, compared with a person (Hirst).
- Keep: verbatim turns, until they are deleted. Content never changes. There is no overwriting and no correction. Two contradictory statements both stay until one of them is forgotten, and retrieval can return both (the design is verified in code; the consequence is inference).
- Forget: whole turns, by age since last retrieval.
- A person loses or changes parts of an episode. MemoryBank drops the whole turn or keeps it exactly.
- Hirst's participants agreed with their own one-week report on .63 of features at 11 months and .57 at 35 months. A MemoryBank replica scores 1.0 on what it kept and has nothing on what it deleted.
- So it diverges from the person in both directions without ever producing a changed version (inference).
- Gist outlives detail, but it is the model's gist. Day summaries survive until their day is empty, and portraits never decay. People also keep central content longer than peripheral detail: in Hirst, crash sites and event order held while airline names declined. Gist supports compact storage (Schacter et al.). Here, though, what lasts is a model-written interpretation, not the person's own (inference).
- The shape of the curve cannot match.
- Between recalls, a single exponential loses the same proportion per unit time. Hirst found change in the first year and then settling.
- Fitted to .63 at 11 months, exp(−t/S) predicts .23 at 35 months, against .57 observed (my computation).
- An exponential matched to the later rate (.57/.63 over 24 months) loses under 5% in the first 11 months, against the 37% people lost (my computation).
- Consistency with a first report is not retention, so this is an analogy, not a fit.
- In this rule, settling can only come from recalls. Its time course would therefore be set by how often retrieval happens to hit an item, not by the person (inference).
- Rehearsal points the other way in the code.
- In Hirst, how much people talked about the event predicted their accuracy on facts, though not their flashbulb consistency.
- Being retrieved is MemoryBank's nearest analogue of being talked about. Under the paper's rule it prolongs retention; under the released code it shortens retention afterwards (above). That is the opposite of the human link (verified in code; effect computed, not run).
- Emotion is where it would diverge most.
- Hirst's participants changed their reports of how they felt more than anything else: .42 and .37 consistency, with new emotions appearing.
- MemoryBank keeps the stored wording of a feeling unchanged. The daily personality analysis records "traits and emotions" and is never forgotten.
- As with perfect retention, emotion is where a MemoryBank replica would diverge most from its person (inference).
- No record of forgetting. Nothing marks a memory as forgotten. After a deletion the replica cannot tell an event that never happened from one it has forgotten, a distinction brief 5.3's ignorance states require (inference).
- Q2: does the portrait converge on generic descriptions?
- The design pushes that way. The global prompt asks for "a highly concise and general summary", and the code adds a response strategy "for the AI lover". The global portrait is re-derived from all the daily analyses, with no input from the user and no decay (verified in code). Each step favours what recurs and is easy to state (inference).
- The paper's examples. The example portraits are positive trait lists that share advice-related terms: "open-minded, curious," and "receptive to advice" (Figure 1), and "seeking advice from others" (Figure 4). These users are talking to an advice-giving companion. The portrait may therefore describe the user's role in the conversation more than the person (guess).
- The released data cannot test this. The simulated users' personalities are given (for example "commanding and unyielding" and "Melancholy and sensitive, kind and gentle"), but no generated portraits were released (verified: those fields are empty).
- For a replica. A self-portrait kept this way is Park et al.'s periodically regenerated identity summary with an explicit instruction to generalise. Expect it to drift toward the base model's picture of a pleasant person and toward the role the replica plays in its conversations (inference).
- A test (proposal). After N days, can a judge pick the right person out of K candidates from the portrait alone? Run it under own, other and empty slots, and compare with the person's own self-description made at the same time.
- Could MemoryBank be a condition in a Q3 test? Yes, as the decay-with-rehearsal arm, reimplemented (proposal).
-
Implement the paper's rule, not the released code: R = exp(−t/S), and S + 1 on each recall.
-
Fix what the paper leaves open:
- the time unit, fitted to the person's retention at the first delay;
- what a recall is: retrieved, or used in the answer;
- whether forgetting deletes or only down-weights. Down-weighting keeps the record. That fits brief 6.2 ("Keep first responses separate from later revisions. Never overwrite."), if the rule is applied to the replica's own store as well (inference).
-
Log every forgetting event. That gives the replica a fourth ignorance state beside brief 5.3's three: forgotten by design.
-
Arms, each under own, other and empty slots:
- verbatim (full context, or EM-LLM);
- decay with rehearsal (this rule);
- extract and overwrite (Mem0);
- a rule fitted to the person's measured forgetting, per kind of content.
-
Target: the person's own reports at two or more delays, scored as in the Hirst summary:
- consistency with the first report;
- transitions between delays (repeated, corrected, changed);
- emotion scored separately;
- confidence.
-
Prediction for this arm (guess):
- exact agreement on what it keeps, and no changed versions;
- lost episodes where the person keeps a changed version;
- a time course set by retrieval traffic.
Fitting the time unit at delay 1 and predicting delay 2 is the direct test of the exponential against Hirst's settling.
-
Cross-references
summaries/carry_on/chhikara2025_mem0.md: the other explicit update policy (extract facts, then add, update or delete). Its LOCOMO table lists MemoryBank at an F1 of 5–10, numbers taken from earlier work, not rerun.summaries/carry_on/fountas2024_em_llm.md: the perfect-retention end of question 3.summaries/carry_on/park2023_generative_agents.md: exponential recency decay used as a retrieval weight (0.995 per game hour), not deletion; the periodically regenerated identity summary.summaries/carry_on/hirst2009_911_memory.md: human consistency, settling and correction over three years.summaries/carry_on/schacter2011_adaptive_distortion.md,summaries/carry_on/nichols2019_false_memory_tasks.md: what people keep, lose and distort, by kind.summaries/carry_on/venkit2026_companion_drift.md: three memory settings at chance on user-state changes.summaries/carry_on/peng2025_funhouse.md: insufficient individuation of twins, the question-2 analogue for portraits.summaries/memory/wu2025_longmemeval.md: knowledge-update and abstention probes.- Brief v1.5, sections 5.3 (provenance fields and ignorance states) and 6.2 (design rules).