ICLR 2025: the PDF header says "Published as a conference paper at ICLR 2025", and arXiv's journal reference reads
"Proc. International Conference on Learning Representations (ICLR), 2025". Read as arXiv 2407.09450 v3 (10 October
2025; v1 12 July 2024 and v2 25 October 2024, not read), cs.AI (cross-lists cs.CL, cs.LG, q-bio.NC), CC BY 4.0. v3
postdates the conference and was not compared with the proceedings version. Huawei Noah's Ark Lab (London) and the UCL
AI Centre. The author list on the abs page matches the PDF: Zafeirios Fountas, Martin A Benfeghoul, Adnan Oomerjee (the
latter two equal contribution), Fenia Christopoulou, Gerasimos Lampouras, Haitham Bou-Ammar, Jun Wang. Code:
github.com/em-llm/EM-LLM-model (not read). Provenance: papers/carry_on/fountas2024_em_llm.provenance.json.
What was read
All 2,503 lines of the pdftotext -layout text of the 39-page PDF: sections 1–6, the references, and appendices A–F
with Tables 1–12, Algorithm 1 and the proof sketch in F.1. The appendices cover the per-task tables against InfLLM and
RAG, the human data, complexity and hardware, hyper-parameters and ablations, further discussion, and proofs. Figures
are images; captions and extracted labels were read.
- Figures 4 and 6–8 (human comparison) give some bar values in the text, but which bar belongs to which method cannot be recovered, except "Random: 2.06" in Figure 4B. It is also unclear which of the two extracted plot blocks belongs to Figures 6, 7 and 8.
- Figures 1 (bottom), 5 and 9–14 carry no usable values in the text.
- The code was not read.
Question
Can a pre-trained LLM handle practically unlimited context, without finetuning, by organising past tokens into events the way human episodic memory is thought to? That means cutting experience at surprising moments and recalling it by similarity and temporal contiguity. And does the resulting segmentation resemble human event segmentation?
Method
-
Context layout. Three groups:
- 128 initial tokens, kept as attention sinks;
- the evicted tokens, which form the episodic memory;
- the local context, with full attention.
Retrieved tokens get a fixed position embedding.
-
Memory formation.
- Surprise is the negative log-likelihood of the actual token, −log P(x_t | x₁…x_t₋₁). A token is a candidate boundary if its surprise exceeds T = μ + γσ over a moving window τ (Eq. 1).
- Boundary refinement. Between consecutive candidates, the boundary is moved to maximise modularity (or minimise conductance). The graph behind these scores has as adjacency matrix the dot-product similarity between one attention head's keys (Eq. 2–4, Algorithm 1). It is a single pass with only rightward moves, which the authors liken to one pass of phase 1 of the Louvain method, started from the surprise segmentation (E.3). Complexity is O(nm) for chunk size m.
-
Retrieval, separately in every layer.
- The similarity buffer holds k_s events found by k-NN between the current query and each event's representative tokens, as in InfLLM.
- The contiguity buffer is a queue of size k_c holding the ±1 neighbouring events of each retrieved event.
- The contiguity ratio k_c/k is 0.3.
-
Hyper-parameters, chosen on LongBench. γ was chosen per model from {1, 2, 3}: 1, 2, 1, 1, 1 for Mistral, LLaMA-3, LLaMA-3.1, Phi-3 and Phi-3.5. The contiguity ratio was chosen from {0.3, 0.5, 0.7} with Mistral. Local + retrieved tokens follow InfLLM: 4K+2K (Mistral), 4K+4K (LLaMA-3/3.1), 1K+3K (Phi).
-
Variants. S (surprise only), SM (plus modularity refinement), S+C (plus contiguity buffer), SM+C (both).
-
Evaluation.
- LongBench (15 tasks; 12K ± 10K tokens per example) and ∞-Bench (6 tasks; > 100K tokens), against InfLLM with fixed-size blocks.
- Against RAG (NV-Embed-v2 or all-mpnet-base-v2 retriever; 300-word chunks, top 5) and against full context, with LLaMA-3/3.1-8B.
- Passkey retrieval up to 10.2M tokens.
-
Human comparison.
- Data: Kumar et al. (2023) podcasts, where human boundaries are a Gaussian-smoothed average across participants. The human boundaries were taken as the most likely positions, "as many ... as our initial surprise-based event segmentation had identified".
- Measures: key-similarity metrics, and the Wasserstein distance between boundary distributions (mixtures of Gaussians).
- Segmentation quality was also measured on PG-19 books, with γ = 10⁻³.
-
Hardware. Nodes of 4 GPUs with 32 GB each and at least 100 GB of CPU memory, "except for the full-context results for which we used an API".
Results
-
Against InfLLM (Table 1, one variant shown per model; LongBench averages):
Model EM-LLM InfLLM Mistral 43.7 41.9 LLaMA-3 47.2 47 LLaMA-3.1 51.3 51.1 Phi-3 35.4 34.5 Phi-3.5 34.9 34.2 - ∞-Bench averages (Tables 3–5): Mistral 66.16 (SM+C) vs 65.78; LLaMA-3.1 65.23–66.69 vs 64.00.
- With LLaMA-3, all variants score below InfLLM on ∞-Bench: 48.81–49.00 vs 50.31. Math.Find loses 27.68%.
- The largest relative gains sit on tiny baselines: PassageRetrieval +40% (Phi-3, 7.50 → 10.50) and Musique +29.7% (Phi-3, 15.05 → 19.52).
- Significance (appendix A.1): p < 0.05 at benchmark level by a two-tailed z-test, "except LongBench with Phi-3.5 with p = 0.23". The authors add: "this isn't the case in the majority of individual tasks".
-
Against RAG and full context.
- LLaMA-3.1 (Table 9), LongBench: EM-LLM_S 51.58, RAG with NV-Embed-v2 36.44, full context 39.30. ∞-Bench: 66.66, 58.97, 66.33.
- RAG with surprise-based chunks scored 25.89.
- LLaMA-3 (Table 8): EM-LLM_SM 47.24, NV-Embed-v2 RAG 39.27, all-mpnet-base-v2 RAG 30.75.
-
Scale. 100% passkey retrieval up to 10.2M tokens.
-
Segmentation quality (PG-19, Table 2; differences from random segmentation).
- Surprise with refinement is best. LLaMA-3-8B: modularity (×10⁵) SM 27.0 ± 35.6, S 13.1 ± 21.5, fixed −1.6 ± 3.6; conductance SC −33.9 ± 9.6.
- Refinement applied to fixed blocks (FM, FC) stays below the surprise-based counterparts.
-
Human data (Figure 4).
- Human-perceived events score higher on the key-similarity metrics than fixed or random events.
- Surprise-only segmentation "achieves very similar results to humans". Surprise-based methods lie closest to human boundaries by Wasserstein distance.
- InfLLM's fixed blocks are worse than random; random scores 2.06 (×10).
-
Ablations.
- Refinement is best in 60% of tasks, contiguity in 44%.
- Smaller γ is better.
- Contiguity ratio 0.3 is slightly preferred.
- More retrieved tokens help only the longest summarisation examples: QMSum 23.24 → 24.47 from 1K to 4K retrieved tokens, 24.30 at 6K (Table 12).
-
Costs.
- Time per 512-token chunk (Mistral-7B): 1.40 s for InfLLM, 1.57 s for S (× 1.12), 2.27 s with refinement (× 1.62).
- KV cache: 9.8 GB at 20K tokens and 488.3 GB at 1M, for 32 layers × 32 KV heads in half precision. It is offloaded to CPU and disk, at about 2 GB of CPU memory per instance.
-
The authors' own qualification (appendix E.1): "we lack explicit empirical evidence that this is how the resulting model makes use of the architecture in order to achieve the results presented, and hence clarify that we only claim an EM-inspired approach, rather than an actual human-like EM process". They list the differences from human memory:
- non-parametric storage;
- no nested event hierarchy;
- no cross-modal integration;
- no consolidation.
Limits
- Headline claims overstate the tables. Each item below was recomputed from the printed tables.
-
"80% of individual task groups". Table 1 has 21 wins, 3 ties and 6 losses in 30 LongBench groups. That is 70% strictly, and 80% only if ties count.
-
"consistently outperforming ... InfLLM across various baseline LLMs". On ∞-Bench with LLaMA-3 every variant is below InfLLM. Appendix A.1 names only Phi-3.5 on LongBench as the exception.
-
The "Max Imp." on each table's average row is the mean of the per-task maxima over the four variants: the best variant, picked per task after the fact. The best single variant improves the average by less:
Model Printed "Max Imp." Best single variant Mistral 4.50% 4.32% LLaMA-3 1.62% 0.62% LLaMA-3.1 1.89% 1.04% Phi-3 8.98% 2.90% Phi-3.5 5.06% 1.87% -
Appendix E.2 claims retrieval +16.6% and multi-document QA +6.4% for S, and +19.4% and +9.2% for SM+C. Only partly reproducible:
- 6.4% matches.
- 16.6% is close to PassageRetrieval alone (16.46%).
- The SM+C column gives 18.09% and 4.55%. The 19.4% equals the mean of the best-variant maxima (19.37%), and 9.2% matches nothing tried (the best-variant mean is 8.98%).
-
"exceeding the performance of NV-Embed-v2 by 30.5% on LongBench and by 11.5% on ∞-Bench". Table 9 gives +41.5% and +13.0% relative to RAG. The 11.5% equals the gap divided by EM-LLM's own score; 30.5% is not reproduced either way (that calculation gives 29.4%).
-
Table 1's variants. Table 1 is said to show "the best single method in terms of overall performance". For LLaMA-3, LLaMA-3.1 and Phi-3 the variant shown is not the one with the best average in Tables 4–6. The differences are ≤ 0.3 points.
-
- The baselines are weak where EM-LLM's lead comes from.
- The full-context baseline ran through an API (appendix C.3) and scores implausibly low on TREC (4.50), SAMSum (8.68), LCC (19.30) and RepoBench-P (18.33).
- Without these four tasks, full context beats EM-LLM on the other 11 LongBench tasks: mean 48.97 against 47.91, recomputed.
- Over all 21 tasks, EM-LLM wins 12, full context 8, with one tie.
- RAG is also very low on few-shot and code tasks, where retrieving five 300-word chunks is a poor fit.
- I read the large margins over RAG and full context as mostly a protocol artefact (inference).
- Tuned on the test benchmark. γ and the contiguity ratio were selected on LongBench, the benchmark reported.
- The human comparison is thin.
- The sources disagree on how many podcasts were used. The text cites Kumar et al.'s three podcasts; the appendix figures show two ("Monkey" and "Tunnel"); Figure 4 speaks of "a human-annotated audio dataset".
- Boundaries are consensus boundaries, not individuals'.
- The number of human boundaries was forced to equal the surprise method's.
- Wasserstein distance was chosen because "standard correlation or discrete distance metrics ... showed very little differences between methods".
- The abstract's "strong correlations" between EM-LLM's segmentation and human events is not backed by any correlation coefficient in the paper. The distances carry no uncertainty.
- Table 2 statistics.
- The standard deviation exceeds the mean in every modularity cell (18 of 18).
- Two deviations are implausible: LLaMA2-7B intra/inter-similarity, FM ± 184.7 and S ± 880.0.
- No test is reported, yet appendix E.1 calls the key similarity within surprise segments "significantly higher".
- Table 2 uses γ = 10⁻³, while the benchmarked system uses γ = 1 or 2. This is unexplained.
- Terminology. The paper calls the token's negative log-likelihood "Bayesian surprise"; Figure 11 labels it "SURPRISAL". Bayesian surprise usually means a divergence between prior and posterior beliefs (my note). The distinction matters when comparing with the human literature.
- The F.1 "proof". It shows that k-NN retrieval approximates softmax attention only by assuming that the k nearest keys hold nearly all the attention weight (α ≈ 1).
- Scope. 7–8B and Phi-mini models only, all with full attention; text only; no hierarchy and no consolidation.
What it means for Kurisutina
- What it keeps: everything.
- Every token's keys and values are stored. Nothing decays, is compressed, merged or revised. Segmentation only groups the tokens, and retrieval chooses per layer and per step what is attended.
- This is the "perfect retention" end of question 3: the design that Schacter et al. and Nichols & Loftus predict will diverge from a person. It keeps detail the person lost, and it has no gist extraction, schema completion or reinterpretation.
- EM-LLM is a retrieval mechanism, not a model of what a person keeps.
- How strongly: there is no strength.
- Surprise sets where events begin, not how well they are kept. Park et al.'s memory stream at least scores importance.
- Access depends only on key similarity to the current query plus contiguity, so old and new events compete on equal terms.
- The local window supplies recency, and the 128 initial tokens are always attended, a structural primacy (my reading).
- There is no forgetting curve.
- How it segments: where the model's prediction fails.
- With a person slot in context, the predictions, and so the boundaries, are those of the slot-conditioned model.
- This gives a question-3 measurement that needs no EM-LLM code. Take the per-token log-probabilities of the same text under own, other and empty slots, and place boundaries by Eq. 1.
- Does the slot move the boundaries toward this person's? Prompt-token log-probabilities are what llama.cpp's
perplexity computation produces. Whether llama-server v0.5.0 returns them for prompt tokens is not checked;
docs/research/local_inference.mdlists prompt log-probabilities for vLLM.
- Compared with human event segmentation: consensus only.
- The evidence is two or three podcasts, count-matched boundaries and a distance ranking.
- The paper's own suggested test was not run (Discussion, after Michelmann et al. 2023b): are model boundaries closer to the human consensus than individual humans are?
- For a replica the question is individual: people segment differently, and the replica should segment as its person does.
- Loaded memories and new experience are not distinguished.
- A loaded slot would be segmented and stored like any experience and retrieved by similarity.
- There are no provenance tags (brief 5.3), no updating of old memories by new ones and no consolidation.
- Interaction between loaded memories and new experience therefore reduces to competition for retrieval.
- Fit with the local models: not a drop-in option.
- EM-LLM patches the attention modules of Hugging Face Transformers and was tested only on full-attention models.
- Qwen3.5/3.8 have full attention in 8 of 32 layers. The other 24 are Gated DeltaNet layers with a fixed-size
recurrent state instead of a KV cache (
docs/research/local_inference.md). - EM-LLM could manage at most the attention layers. The DeltaNet state is itself a lossy memory of the whole stream, so the hybrid base model already combines forgetting and exact memory (inference). Porting is untested; llama.cpp has no such mode.
- Size: at the doc's figure of about 32 KB of KV per token for Qwen3.5-9B, 1M tokens of experience is about 32 GB, on CPU or disk as in appendix C.3 (estimate).
- How to test it against a person (proposals, in order of cost).
-
Segmentation. The person marks event boundaries in a new recording or story, at a fine and a coarse grain. Compare the model's surprise boundaries under own, other and empty slots with the person's own boundaries, and with a leave-one-out consensus of other people. Score agreement within a tolerance window, plus the paper's Wasserstein distance. Check the method first on Kumar et al.'s data, which the paper describes as released, using a local model.
-
Recall structure. Person and replica freely recall the same material after delays. Score:
- which events are recalled;
- the order, as a lag-conditional response probability (Howard & Kahana, the paper's Figure 3A);
- whether boundary content is favoured, since boundaries are "access points" (Michelmann et al. 2023a, cited);
- how recall decays.
EM-LLM predicts flat retention, and a contiguity effect whose size is set by the contiguity ratio, not by the person.
-
Divergence in both directions, as in the question-3 plan: detail the replica keeps that the person lost, and updates the person registered that the replica did not.
-
Misinformation. In EM-LLM it is stored as a separate event. Test whether the replica keeps it apart or merges it with the original, and compare that with the person.
-
- Design lesson. Segmenting by surprise is a plausible and cheap way to cut experience into episodes, and it can be made person-specific through the slot. What is kept and how strongly cannot be left at "everything". It has to be set from the person's own forgetting and distortion, measured per kind (Nichols & Loftus).
Cross-references
summaries/carry_on/park2023_generative_agents.md: the memory stream (recency, importance, relevance; reflection), the contrasting agent-memory design.summaries/carry_on/venkit2026_companion_drift.md: current memory settings fail on user-state changes (chance) and temporal order (below 0.5).summaries/carry_on/schacter2011_adaptive_distortion.mdandsummaries/carry_on/nichols2019_false_memory_tasks.md: the human side of what is kept, lost and distorted.summaries/carry_on/microverse2026_identity_drift.md: an agent revising itself from its own memories; the revision rate set by the reflection schedule.summaries/carry_on/lu2026_assistant_axis.md: identity follows the latest message, not accumulated context.docs/research/local_inference.md: the hybrid Gated DeltaNet architecture, KV size per token and runtimes.