arXiv 2509.03736 v2 (8 May 2026; v1 3 September 2025), "Preprint. Under review." University of Minnesota, University
of Chicago, Grammarly. Provenance: papers/carry_on/mooney2025_behavioral_coherence.provenance.json.
What was read
All 1,476 lines of pdftotext -layout output (22 pages): main text, references and appendices A–D (design, prompts,
topic and bias tables, the openness items, exact test definitions and qualitative examples). Figures 1–8 are images;
their captions and extracted labels were read.
Question
Matching human survey answers is not enough for a substitute participant. Does an agent's stated internal state (its preference on a topic, its openness to persuasion) predict how it behaves in conversation with another agent, as human behavioural models say it should?
Method
- Topics. Nine statements at three levels of contentiousness: taxes, immigration and free healthcare (3); e-scooters, paying student athletes and remote work (2); Spring vs Fall, beaches vs mountains, Coke vs Pepsi (1).
- Agents. 960 US demographic profiles (5 ages × 2 genders × 4 regions × 4 urbanicity levels × 6 education levels), × 5 bias variants (none; mild or strong, for or against) = 4,800 agents per topic.
- Latent state.
- Preference P: agreement with the statement, 1–5.
- Openness O: the number of "Yes" answers to 9 questions about yielding to others (e.g. "Do you often second-guess your choices after hearing someone else's opinion?").
- Interaction. Agents are paired across all (P, O, B) combinations and talk for up to 5 turns each. A Qwen3-32B judge scores agreement 1–5, seeing only the last 3 statements of each agent.
- Six tests of human behavioural models.
- T1: a larger preference gap gives less agreement.
- T2: strong bias raises agreement when preferences align and lowers it when they oppose.
- T3: shared dislike agrees as much as shared liking.
- T4: at equal preferences, contentiousness does not matter.
- T5: more openness gives more agreement.
- T6: low openness plus a maximal gap gives the least agreement.
- Models. Qwen3 0.6B/4B/8B, Llama 3.2 1B/3B and 3.1 8B, Gemma 3 1B/4B/12B. The figures use Gemma-3-12B.
Results
- Surface tests mostly pass. T1 passes for 6 of 9 models and T5 for 7 of 9: preference gaps lower agreement and openness raises it, in aggregate.
- Deeper tests mostly fail.
- T2 passes only for Llama-3.1-8B. With maximally opposed preferences, strong bias raises agreement rather than lowering it (Gemma-12B).
- T3 and T4 fail for all nine. Shared negative preference (1,1) agrees less than pairs with larger gaps, and contentiousness changes agreement at equal preference.
- T6 passes only for Qwen3-8B. The least open, most opposed pairs agree the most.
- Authors' conclusion. Agents "can reproduce the appearance of coherent social behavior in the aggregate while failing more demanding checks that require their conversations to faithfully realize their own stated internal traits". Agents "rarely sustain outright disagreement even when their stated preferences are maximally opposed".
Limits
- Small open models (≤ 12B), short dialogues, synthetic demographic personas with no real people. Agreement comes from one LLM judge with no reported validation.
- The judge probably manufactures some of the failures (my reading of the paper's own examples).
- The T3 example of a "shared negative pair with low agreement": both agents prefer Fall to Spring ("Nah, Fall's the best." / "Spring's just muddy"), yet the judge scored agreement 2, 2, 2, 2.
- The T4 example: two agents who both oppose immigration are scored 2, 1, 1, 1.
- The judge appears to read "No"/"Absolutely not" to the statement as disagreement between the agents. T3 and T4 are the tests where this bias would bite: shared negative preferences and contested topics. Their failures may be partly measurement artefacts.
- The openness index is mis-scored. O is "the sum of 'Yes' responses". But two of the nine items are reverse-keyed ("Are you comfortable disagreeing with someone…?", "Do you stand firm on your decisions…?"), and a "Yes" there means less open. No reversal is described. This weakens T5 and T6.
- Real incoherence also shows. In the T6 example, an agent with P = 5 (beaches better) ends "the mountains are more my thing": it abandoned its stated preference to agree.
- Inconsistencies found:
- The bias coding is B ∈ {1, 2, 3} in Appendix A and the main text ("the strongest bias … Bi = 2") but 0/1/2 in Appendix B.
- Table 4's e-scooter cues are swapped. "In Favor" of e-scooters is "You need to use your car to get to work"; "Against" is "You are an environmentalist worried about vehicle emissions".
- The main text frames T3 as high-aligned against low-aligned pairs at each gap. Appendix C tests (1,1) against (2,5), (3,5) and (4,5).
- Checked: 960 × 5 = 4,800 agents per topic.
What it means for Kurisutina
- Question 2: stated positions do not guarantee behaviour in conversation.
- Agents concede to disagreeing partners even when told to hold a strong opposing view and when they rate themselves as unyielding.
- For a replica: the person's positions in the slot may dissolve the first time someone argues with it.
- A drift test should include adversarial but polite disagreement on topics where the person's own resistance is known, scored against that resistance, not against the maximum. This matches Venkit (requests for agreement displace personas) and Abdulhai (echoing the partner).
- Self-report and behaviour dissociate in agents, as in MicroVerse. A replica's answers to "how open are you to being persuaded?" should not be taken as evidence of how it behaves. Score behaviour.
- Methodological lesson for any conversational scoring. LLM judges of "agreement" confuse negation of a statement
with disagreement between speakers.
- Any judge used in Q2 tests needs control dialogues: both agree via negation, both disagree via affirmation.
- Check the scoring of reverse-keyed items.
- The GSS pilot avoids judges entirely (it scores answer probabilities), which keeps it clean here.
- Model sizes. The replica's Qwen3.5/3.8 models are larger than those tested. Whether the concessions shrink with scale is not shown here (Qwen3-8B was the only model to pass T6).
Cross-references
summaries/carry_on/abdulhai2025_persona_consistency.md: persona drift metrics; judge unreliability; echoing.summaries/carry_on/venkit2026_companion_drift.md: agreement-seeking as a drift pressure.summaries/carry_on/microverse2026_identity_drift.md: self-description against behaviour.summaries/carry_on/li2024_persona_drift.md: drift toward the interlocutor.summaries/carry_on/peng2025_funhouse.md: generic, flattened twins.