UIST '23 (ACM), DOI 10.1145/3586183.3606763; arXiv 2304.03442 v2 (6 August 2023). Stanford, Google Research, Google
DeepMind. Code github.com/joonspk-research/generative_agents (not read).
Provenance: papers/carry_on/park2023_generative_agents.provenance.json.
What was read
All 22 pages of the extracted text: sections 1–9, the 109 references, appendix A (architecture optimisations) and appendix B (all 25 interview questions with Klaus Mueller's sample answers). Figures are images; captions were read.
Question
What architecture lets an LLM-driven agent act consistently with its accumulated experience over time, remember, plan and form relationships, rather than respond only to the current moment?
Architecture
- Memory stream. Every perception, as a natural-language record with a creation time and a last-access time.
- Retrieval. Score = recency + importance + relevance, each min–max normalised, all weights 1.
- Recency decays exponentially: factor 0.995 per game hour since last access.
- Importance is the LLM's 1–10 rating of "poignancy" when the record is made (cleaning a room 2, asking a crush out 8).
- Relevance is embedding cosine similarity to the query.
- The top records that fit in context go into the prompt.
- Reflection. Runs when the summed importance of recent events passes 150 (two to three times a game day).
- The 100 latest records produce "3 most salient high-level questions".
- Retrieval for each question, then "5 high-level insights … (because of 1, 5, 3)".
- The insights are stored as reflections with pointers to their evidence; reflections can reflect on reflections.
- Planning. A broad day plan (5–8 chunks) from the agent summary and the previous day, decomposed into hours and then 5–15-minute steps. Plans are stored in the stream. At each step, "Should X react?" can trigger re-planning; dialogue is conditioned on each agent's retrieved memories of the other.
- Identity summary (appendix A). The "[Agent's Summary Description]" used in almost every prompt is regenerated periodically by retrieving on "[name]'s core characteristics", "current daily occupation" and "feeling about his recent progress in life", then summarising.
- Setting. 25 agents in "Smallville" (Phaser sandbox), each seeded with one paragraph whose semicolon-separated phrases become initial memories; gpt-3.5-turbo. The world is a tree of areas and objects; each agent keeps its own possibly outdated subtree.
Evaluation and results
- Controlled interview. Agents after two game days answered 25 questions (self-knowledge, memory, plans, reactions, reflections). 100 Prolific raters ranked five conditions by believability: TrueSkill full architecture 29.89, no reflection 26.88, no reflection or planning 25.64, crowdworker writing in the agent's voice 22.95, no memory at all 21.21 (d = 8.16 full vs none; Kruskal–Wallis H(4) = 150.29; all pairs different except crowdworker vs no-memory). Ablations received the same accumulated memories, so the authors call the differences conservative.
- End to end (two game days). Knowledge of Sam's candidacy spread from 1 to 8 agents (32%) and of the party from 1 to 13 (52%), with no hallucinated knowledge of either. Network density rose from 0.167 to 0.74; 6 of 453 answers about knowing another agent (1.3%) were hallucinated. Of 12 invited, 5 came to the party; 3 cited conflicts, and 4 said they were interested but made no plan.
- Failure modes named by the authors.
- Retrieval misses and partial fragments: Tom remembers planning to discuss the election at a party he is not sure exists.
- Embellishment: Isabella adds that Sam will "make an announcement tomorrow".
- Base-model knowledge leaking into memory: a neighbour called Adam Smith "authored Wealth of Nations".
- Worse place choices as memory grows.
- Physical norms not conveyed in text (one-person bathroom, shops closed after 5 pm).
- Instruction-tuned politeness and over-cooperation. Isabella rarely said no to other agents' suggestions, and "over time, the interests of others shaped her own interests" (she became "very interested in literature").
- Risks the authors name. Parasocial attachment (disclose the agent's nature; do not reciprocate confessions of love), errors in inference about users, deepfakes and tailored persuasion (keep audit logs), over-reliance replacing human input, and "memory hacking": a crafted conversation convincing an agent that a past event occurred. Cost: thousands of dollars in tokens and several days of computation for two game days of 25 agents.
Limits
- The target is believability to observers, judged by crowd raters, not fidelity to any real person. The agents are fictional, seeded with one paragraph.
- Two game days; no long-horizon stability test; one 2023 model.
- The interview evaluation freezes agents at one moment; ablations did not live their own histories.
What it means for Kurisutina
- This is the reference design for the replica's own memory formation (question 3), and it has none of the needed controls. Everything is kept, and importance is judged by the base model's sense of poignancy, not the person's. Reflections are the base model's inferences. The summary of "who I am" is re-derived from memories on a schedule. Each of these lets the base model, not the person, decide what the replica becomes.
- Identity erosion is already visible over two game days. Isabella's interests were reshaped by other agents' suggestions because the instruction-tuned model rarely refuses. For a replica, the equivalent is taking on the interests of whoever it talks to. That is the user's drift criterion failing in a documented case, and it fits Li et al.'s agent adopting its interlocutor's instruction.
- What a replica would change.
- Keep the loaded self (the person's slot) as an anchor, not something re-summarised from new memories.
- Let the person's own salience and forgetting, measured, decide what is kept, instead of the model's poignancy score.
- Tag every reflection as the replica's inference, separate from the person's reported memories (brief 5.3's provenance fields).
- Score it on fidelity to the person's changes, not believability.
- Two failure modes to test for directly.
- Base-model knowledge entering memories (the Adam Smith error), the same hyper-knowledge as in Peng et al.
- Memory hacking, which for Amadeus is suggestion: the replica accepting a false past because someone asserted it.
The human memory literature already in
summaries/memory/(misleading suggestions, Lindsay 1990; source monitoring) gives the human rate to compare with.
Cross-references
- Park et al. 2026 (brief ref. 12,
summaries/12_park2026.md): the same group's later agents grounded in real people's interviews, static rather than living. - Li et al. 2024 (
summaries/carry_on/li2024_persona_drift.md) and Lu et al. 2026 (summaries/carry_on/lu2026_assistant_axis.md): mechanisms for the drift seen here. summaries/memory/lindsay1989_eyewitness_source.md,summaries/memory/johnson1993_source_monitoring.md,summaries/memory/lindsay1990_source_suggestions.md,summaries/memory/roediger1995_false_memories.md: the human side of memory implantation and source confusion.