arXiv 2609.19913 v1 (17 September 2026), cs.LG. Télécom Paris (Institut Polytechnique de Paris), University of
Salerno, University of Kentucky. Preprint, no peer review indicated. Provenance:
papers/carry_on/berjawi2026_opinion_twins.provenance.json.
What was read
All 966 lines of the pdftotext -layout text of the 15-page PDF: sections 1–9, Tables 1–4, Algorithm 1, the acronym
appendix and the references. Figures 1–5 are images; their captions and in-text descriptions were read.
Question
Can LLM agents, cloned from real Twitter users and placed in the real interaction network, reproduce how those users' opinions evolve over time, better than classical opinion-dynamics models?
Design
- Data. Two Kaggle Twitter datasets: COVID-19, over 1.1 million tweets; US election 2020, over 1.7 million.
- Each user's daily "opinion" in [−1, 1] is the sentiment polarity of their tweets (RoBERTa), not an annotated stance.
- The network is built from mentions, with weights counted over the full observation window. The largest connected component is kept: about 700 agents (COVID) and about 500 (election).
- Agents. Mistral-7B via LangChain.
- Each agent gets: a persona vector (ethos/pathos/logos style from a persuasion classifier); centrality measures; activity counts; sentiment and emotions; stubbornness λ (OLS on the calibration period); and influence β (from the Friedkin–Johnsen matrix).
- At each step it receives its own opinion, its full memory of past opinions and reasoning text, and its neighbours' simulated opinions with their last calibration-period post. It returns a new opinion, and the first number in the output is taken.
- Protocol. The first 70% of time is calibration and the last 30% validation. The twin starts from the real state at the boundary and then runs autonomously, with no further real data. 15 runs per temperature.
- Baselines. DeGroot, Friedkin–Johnsen, Hegselmann–Krause and Deffuant–Weisbuch, each with a small parameter grid.
- Metrics. Per-agent MAE, and absolute discrepancies in homophily r, variance, and EMD.
- Model selection. For every model, including Mistral's temperature, the configuration was chosen "after evaluating all four metrics … over the validation horizon".
Results (as reported)
- Mistral-7B, τ = 0.8. MAE 0.150 ± 0.02 (COVID) and 0.121 ± 0.02 (election), against 0.309 and 0.312 for the best
classical model (Hegselmann–Krause): "more than 50%" lower.
- Δr 0.120 and 0.180; ΔVar 0.106 and 0.115.
- ΔEMD 0.293 ± 0.10 and 0.311 ± 0.15, where its advantage is least consistent.
- Ablations (MAE COVID / election): full 0.150 / 0.121; without attributes 0.380 / 0.299; without memory 0.265 / 0.216; without social exposure 0.255 / 0.195.
- Temperature. Mistral is stable across temperatures 0–1 (MAE 0.150–0.214 COVID), while the classical models are sensitive to their parameters.
- Mean simulated trajectories are said to track the real series, "capturing gradual trends and short-term fluctuations alike".
Limits
The authors' own:
- sentiment is a proxy for stance;
- two Twitter datasets;
- one LLM;
- a static network;
- computation cost.
Mine, from the text:
- The headline is implausible as reported (my inference).
- The only constant-prediction baseline available is Friedkin–Johnsen with α = 0. It holds every agent at its calibration-derived intrinsic opinion, and has MAE 0.512 (COVID) and 0.489 (election) in Table 4.
- An agent that runs autonomously, with no validation-period data, can reach MAE 0.15 against daily sentiment only if it predicts the real day-to-day swings. The paper claims it tracks "short-term fluctuations".
- The text gives no mechanism by which it could know them. Nor does it report a persistence or per-user-mean baseline, or compute MAE on a stated aggregate.
- Without the code the result cannot be checked, and it should not be relied on.
- Selection on the test period. All configurations, including Mistral's temperature, were selected on the validation horizon.
- Leakage of structure. Network weights use mention counts over the full window, including the validation period (377 and 290 validation edges).
- No person-specific controls. There is no empty-attribute arm scored as a baseline beyond the ablation, and no other-agent's-attributes arm. "Remove attributes" makes agents homogeneous, which is the empty slot, but nothing tests whether an agent's own attributes beat another's.
- Inconsistencies found:
- Tables label Friedkin–Johnsen's parameter λ, although the text defines its grid over α and stresses that α is not the agents' λ.
- The text cites a "Table V", but only Tables 1–4 exist.
- An equation reference is broken ("Eq. (??)").
What it means for Kurisutina
-
Not usable as evidence that LLM twins predict individuals' opinion change. Opinions are sentiment proxies, the baselines are weak and selected on the test period, and the key comparison lacks an explanation. It is logged so the claim is not cited later from its abstract.
-
A checklist it fails, which the GSS pilot is built to pass:
- a "no change" baseline and a population-change baseline: the pilot has P(later | earlier) and P(later | earlier, event);
- selection frozen before the test period: the pilot's freeze;
- own against other against empty: the pilot's conditions;
- a real outcome measure, the person's later answer, rather than a sentiment proxy.
Any carrying-on result Kurisutina reports should state these four explicitly.
-
One design point worth keeping. Agents that keep their full memory of their own past reasoning (the memory ablation hurt most after attributes) are the open-loop condition of Namazova et al. An evaluation of carrying on needs that condition, scored against the person's real trajectory. The paper shows how easily such a comparison goes wrong when the baselines are weak.
Cross-references
summaries/carry_on/namazova2025_open_loop.md: open-loop generation against teacher-forced prediction.summaries/carry_on/hullman2026_validating.md: held-out, pre-specified evaluation as the condition for validity.summaries/carry_on/gss_mr119.md,summaries/carry_on/hout2016_gss_reliability.md: "no change" is a strong baseline for most answers.docs/research/gss_pilot_design.md: baselines, freeze and conditions.