Kurisutina

HumanLM: Simulating Users with State Alignment Beats Response Imitation

arXiv 2603.03303 v1 (submitted 7 February 2026; the only version). The PDF footer reads "Preprint. March 5, 2026." cs.CL, cross-listed cs.AI. Stanford University, New York University, Accenture. Licence CC BY 4.0 (abs page). Project site humanlm.stanford.edu, code and data not read. Preprint, no peer review indicated. Provenance: papers/carry_on/humanlm2026_state_alignment.provenance.json.

What was read

All 1,422 lines of the pdftotext -layout text of the 27-page PDF: sections 1–8, the impact statement, the references and appendices A–F. The appendices hold the dataset statistics (Table 2), baselines, training details, Tables 3–9, and the profile, judge and system prompts in full. Figures 1, 3 and 8 carry extractable text, which was read. The other figures are images: captions and extracted labels only. The values printed in Figures 6 and 7 were assigned to their bars by word coordinates (pdftotext -bbox). Figures 11–17 are screenshots of the user-study interface: captions only. The arXiv abs page was read for metadata. Not read: code, the Humanual data, the project website.

Question

User simulators trained to imitate real users' replies copy surface style. Does a model that first writes the user's latent "state" (belief, goal, value, stance, emotion, communication), and is rewarded for agreement with the real reply, simulate real users better than imitation does?

Method

  • Task. Data are triples: a user profile, a context (post, article, thread or conversation) and the user's real reply. Eq. 1 treats a reply as a bag of latent states and asks the model to match the real reply's set. A latent state is a short natural-language attribute, such as "empathy towards victims".

  • What "state" is. Six dimensions in four groups: cognitive (belief, goal; after the Belief–Desire–Intention framework), normative (value, stance; positioning theory), affective (emotion) and linguistic (communication). The generation prompts (appendix E.3) define them:

    • belief: a general assumption about people or the world; "Belief is not specific to a target or event";
    • goal: what the user is trying to do with this comment;
    • value: what should matter;
    • stance: agreement toward an explicitly named target;
    • emotion: the emotion, its intensity and its target ("Moderate heartbreak for the wildfire victims");
    • communication: tone and how the message is structured.
  • How a state is inferred and updated. The model writes it from the profile and the context, afresh for every reply. Nothing is carried from one reply to the next, and no update step is described (verified). In the multi-turn Chat data, earlier turns are simply part of the context. At test time the states are not generated separately: the model writes a reasoning trace that should contain them, then the reply (verified, sections 3.3–3.4).

  • Training. Qwen3-8B with GRPO: group size 4, batch size 32, sampling temperature 0.8, at most 1,024 tokens.

    • In each batch one state dimension is sampled. The model generates that state, and a judge (gpt-5-mini) scores the four rollouts together, by comparison, for agreement with the real reply along that dimension.
    • Reply generations in the same mixed batch are rewarded by reply ("response") alignment.
    • The reward is therefore always agreement with the user's real reply, never agreement with the profile.
  • Judge rubric (appendix E.2). The judge extracts 1–3 key points from the real reply along the dimension and scores each match from 0 to 1. Coverage C is the mean match. A penalty P covers extra or conflicting content, including excess length. Score = max(0, min(1, C − P)). The judge sees the context, the real reply and the generations. It never sees the profile.

  • Profiles. For each user, claude-4.5-haiku summarises at most the earliest 20 training-split replies into JSON:

    • demographics only where stated;
    • 8–12 interests, 8–12 values, 8–12 communication-style phrases, each quoting the user;
    • reply-length statistics and frequent words.

    Users need at least 10–20 replies (the threshold varies by dataset) and at most 1,000. Users seen only in validation or test are dropped. Chat has no profiles, "due to a lack of precise user identifiers".

  • Benchmark "Humanual" (Table 2), all from public sources. Totals: 26,179 users, 66,531 contexts, 216,309 replies (my sums; the paper rounds to 26k, 66k and 216k).

    • News: YouTube comments on BBC and CNN videos; 10,900 users, 43,273 comments.
    • Book: Amazon book reviews; 209 users with 192.04 reviews each on average; 1998–2023.
    • Opinion: Reddit r/AITA; 4,567 users, 992 threads, 45,716 replies.
    • Politics: Medium political blogs; 5,303 users, 50,273 replies.
    • Chat: WildChat user turns; 4,801 users, 29,334 replies.
    • Email: Enron; 399 users, 7,577 replies.
  • Splits. Contexts are sorted by time: 90% train, 2% validation, 8% test. Test contexts are therefore later and unseen, for the same users. The main text splits Chat by turns within each conversation instead.

  • Evaluation. Reply alignment and per-dimension state alignment by claude-4.5-haiku, a different family from the training judge. Tables show them on a 0–100 scale (inferred: Figure 5's axis runs 0.22–0.28 where Table 1 gives 25.6). Embedding cosine similarity serves as a judge-free check; the embedding model is not named. Gemini-3-pro re-judges Politics only.

  • Baselines. All are Qwen3-8B and get the same profile and context:

    • the base model, without and with its thinking mode (Qwen3-8b-think);
    • SFT on the real replies, and SFT-think with synthetic thoughts written by gpt-5-mini;
    • GRPO and GRPO-think, rewarded on reply alignment only;
    • UserLM, on Chat only.
  • User study. 111 Amazon Mechanical Turk workers:

    • answered open-ended questions on their values and communication style, which became their profile;
    • wrote their own reply to one of 79 r/AITA test posts;
    • annotated their own reply along the six dimensions (Figure 15);
    • rated three simulated replies (Qwen3-8b-think, GRPO-think, HumanLM, in random order) for similarity to their own and for humanlikeness, on 1–10 scales.

    Pay was $9 per task, 32.1 minutes on average.

Results

  • Benchmark (Table 1, reply alignment, 0–100). HumanLM is best on all six datasets. Averages: HumanLM 13.2, GRPO-think 10.4, GRPO 10.3, Qwen3-8b 9.5, SFT-think 8.6, Qwen3-8b-think 8.4, SFT 6.5.
    • Against the best baseline per dataset: News 9.55 vs 7.92; Book 18.5 vs 13.6; Opinion 25.6 vs 23.8; Politics 12.6 vs 10.9; Chat 6.08 vs 5.83; Email 6.71 vs 5.90.
    • The headline 16.3% is the mean of these six relative gains, each against that dataset's best baseline (recomputed; it matches). In absolute terms the gains are 0.25–4.90 points, mean 1.85 (my computation).
    • Absolute levels are low. The base model's average is "around 10%"; the best cell is 25.6.
  • SFT is worst. Per the authors, SFT replies "mimic user tones well" but are too long and "frequently hold opposite opinions compared to the ground truth". No count is given.
  • State alignment (Figure 4; Tables 4–9). HumanLM has the highest average on five datasets. On Chat, GRPO is higher (11.4 against 10.8). HumanLM trails some baseline on belief, goal and emotion in Book, belief in Opinion, value and stance in Politics, and belief, value and emotion in Chat (checked cell by cell in Tables 5–8).
  • Embedding similarity (Table 3). 45.7 against 43.7 (GRPO-think) and 42.5 (Qwen3-8b-think); the "7.5%" gain over the latter is correct (recomputed). On Opinion HumanLM is below GRPO-think (46.21 against 46.33).
  • Second judge (Figure 6, Politics). Both judges rank Qwen3-8b-think, SFT-think, GRPO-think and HumanLM in the same order. Claude gives 7.0, 9.2, 10.6 and 12.6 (Table 1's Politics column); Gemini 7.2, 8.1, 10.8 and 14.1 (values assigned by coordinates).
  • Training dynamics (Figure 5). HumanLM's checkpoints spread wider on state and reply scores ("23% and 104% higher" spans); the authors describe GRPO-think as "stuck".
  • Which dimensions matter (Figure 7, Opinion, 1k simulated replies). Correlation of each dimension's state score with the reply score: goal 0.66, stance 0.61, communication 0.59, value 0.55, belief 0.47, emotion 0.46 (assigned by coordinates).
  • User study (Figure 9).
    • Mean similarity: HumanLM 6.5, Qwen3-8b-think 6.1, GRPO-think 5.9. Win rates 41.4%, 30.6% and 27.9%.
    • Paired one-sided Wilcoxon tests over 111 participants: p = 0.0279 against Qwen3-8b-think, p = 0.00284 against GRPO-think.
    • Mean humanlikeness: 7.5, 7.4 and 6.9. No test is reported.
    • Participants said HumanLM more often matched their stance and the considerations behind it, and pitched emotional intensity better.

Limits

  • It cannot separate this person from a typical person (verified from the design).
    • Every model gets the same profile. There is no condition with no profile, with another user's profile, or with another real user's reply to the same post. The judge never sees the profile.
    • A model that writes the most typical reply to a post, ignoring the profile, would score well wherever the typical reply is also this user's reply (inferred). On r/AITA that is probably common, because the post largely fixes the verdict (guess).
    • The data would allow the missing control. Opinion has about 46 real replies per thread and News about 7 per video (my computation from Table 2). Contexts differ slightly within a thread, since earlier comments are part of each.
    • The user study has the same gap: all three models get the participant's own profile.
  • The state is a decomposition of one reply, not a carried state. It is regenerated for each reply from a static profile and the context. The profile comes from the earliest replies and is never updated, while the test contexts are the latest. Whether error grows with time since the profile, or with the user's own change, is not analysed.
  • At test time the states are implicit. The claim that reasoning traces contain aligned states rests on three case studies (Figure 8). The paper does not say which generations are scored for state alignment. The baselines write no explicit states, so presumably the replies (inferred).
  • Judges throughout. Rewards and scores come from LLM judges extracting "key points". Robustness is checked with one other judge on one dataset. The judge-free embedding similarity moves little. No confidence intervals, test-set sizes or seeds are given for the benchmark.
  • Small user-study effects. 6.5 against 6.1 on a 10-point scale. Two one-sided tests, uncorrected; with a Bonferroni correction for two comparisons, 0.0279 would not pass 0.025 (my computation). The participants' own six-dimension annotations were collected, but no analysis of them is reported.
  • Chat has no profiles, so its results are about users in general, not individuals (verified, section 4).
  • Inconsistencies found:
    • Section 1 says 55.9% of participants rated HumanLM mostly similar or nearly identical to their own reply (best baseline 45.0%). Section 6 gives 68.6% for the same two categories.
    • The PDF abstract claims the highest humanlikeness. The arXiv abstract says "competitive human-likeness scores". The means are 7.5 against 7.4, untested.
    • "relative improvements of 38% and 17% over base-think" and GRPO-think: Table 1 gives 57% and 27% on the averages, or 90% and 40% as means of per-dataset gains (my computation). 38.9% is the gain over Qwen3-8b without thinking (my computation from the printed averages).
    • "the highest alignment scores on 80% of the latent states" (Figure 4): 18 of 24 cells strictly (75%), 19 counting one tie (my count from Tables 4–7). Over all six datasets and all baselines: 26 of 36 (72%; my count).
    • Chat split: earliest 80% of turns to training (section 4), but a 90/2/8 split by context for each dataset (appendix A).
    • Table 2 lists 4,801 Chat users, yet Chat is said to lack user identifiers.
    • Table 9: Qwen3-8b-think's Email average is printed 17.8; its cells give 17.92 (my computation). All other averages in Tables 1 and 3–9 match to rounding.
    • Table 3: HumanLM's Chat value equals its Opinion value exactly (46.21); possibly a copying error (guess).
    • The conclusion speaks of "26k worldwide user responses"; Table 2 gives 26k users and 216k replies.
    • "$9 per task" at 32.1 minutes is $16.82 an hour, not "approximately $18" (my computation).
    • arXiv lists v1 on 7 February 2026; the PDF footer says 5 March 2026. An identifier in the 2603 range fits announcement in March (guess).

What it means for Kurisutina

Question 1, state and slot:

  • HumanLM's "state" is not what the slot must hold (inferred).
    • Its six dimensions describe the output of an appraisal: where the person stands on this post, what they feel about what, what the comment is for.
    • Our question-1 synthesis says the slot must hold what produces such appraisals: usual reactivity, a slow-changing current state (load), agency, and recent events in order (Sliwinski, Brüning, Kritzler, from their summaries). HumanLM has none of these.
    • Its profile is a static summary of traits and style. Its trait-like dimensions (belief, value) are re-inferred from that profile at every reply. A replica's slot should hold them explicitly, from the interviews, rather than let the model re-derive them each time.
  • The dimensions are a usable vocabulary for scoring appraisal. Stance, emotion with intensity and target, and goal are close to what the proposed held-out appraisal test asks. Rated appraisal scales (ECQ) stay preferable for us, because they need no judge (inferred).
  • The user study contains the design we want, unanalysed. Participants wrote a reply and annotated their own stance, emotion, belief, value, goal and communication. A person's own annotation of their reaction is a cheap target for a replica. For the interviews: collect a few new vignettes with the person's reply and their own ratings of it (proposal).
  • The training reward has the right shape. It rewards agreement with the user's real, later replies, not with the profile. That is the constraint the goal doc drew from Abdulhai et al. But it is one model for all users, it has no notion of change over time, and training is out of scope now (needs rented GPUs; proposal for later).
  • Imitation copies the voice and misses the position. SFT reproduced tone but often took the opposite opinion. If a replica is ever fine-tuned on the person's texts, voice fidelity will not show position fidelity. Judge-free style markers measure only that half of fidelity (inferred).

Question 2, drift measurement:

  • No drift is measured here, and no person-against-average contrast. Any HumanLM-style scoring in our tests needs the controls it lacks:
    1. the empty slot;
    2. another person's slot;
    3. another real person's reply to the same post, scored against this person's reply. This is the ceiling for the average person with similar views, the text analogue of GSS baseline 3 (same earlier answer, same event).
  • Per-dimension scoring as a diagnostic. Scoring a replica's reply against the person's real reply, dimension by dimension, shows where it drifts: right stance, wrong emotion, the base model's communication style. The judge takes its key points from the person's real reply, so the anchor is the person, not the slot text. It is still an LLM judge and needs the control items and held-out judge the goal doc asks for (Abdulhai, Mooney).
  • It complements the Persona-Retention projection; it does not replace it. PR (anchored on the person's answers and the empty slot) measures movement toward the base model on sealed probes. HumanLM-style scores measure agreement of free replies with the person. Neither tells us drift direction on free replies; that needs the same free reply scored under own, other and empty slots (proposal).
  • A carrying-on data set sits in the benchmark, unused. Humanual-Book has 209 users with about 192 reviews each, spanning 1998–2023 across the dataset. Profiles come from the earliest reviews and tests from the latest. It could test whether a replica built early predicts a person's later judgements better than a population model, and whether error grows with elapsed time (proposal; availability and licence of the data not checked).
  • Expressed, not private. Every Humanual target is a public reply written for an audience (commenters, readers, a chatbot, colleagues). People tune what they say to an audience strongly and carry little of it over (Wagner et al., from its summary). A score against public replies therefore measures expressed accommodation. The carried-over part needs private probes before and after the conversation, as the goal doc proposes (inferred). Our interviews have an audience too: the interviewer.
  • Judge-free measures. The one judge-free measure here, embedding similarity, barely separates the methods (42.5–45.7 across three models). Coarse embeddings will not show drift. Specific markers compared with the person's own texts (Breithaupt et al.'s word classes) remain the better judge-free option (inferred).
  • As a test condition. State first, then reply could be a replica condition. But HumanLM's gain comes from RL: prompted reasoning alone (Qwen3-8b-think) scored below the plain base model on average (8.4 against 9.5). So prompt-only state-first generation is not shown to help (inferred).

Cross-references

  • summaries/carry_on/abdulhai2025_persona_consistency.md: RL for persona consistency; why a fidelity reward must be agreement with the person's later answers; judge failures.
  • summaries/carry_on/venkit2026_companion_drift.md: Persona Retention with the empty-slot anchor.
  • summaries/carry_on/peng2025_funhouse.md: twins that regress to a generic persona, the failure this benchmark cannot detect.
  • summaries/carry_on/mooney2025_behavioral_coherence.md: stated state against behaviour in conversation.
  • summaries/carry_on/kritzler2022_event_perception_profiles.md, summaries/carry_on/haehner2022_event_perception.md, summaries/carry_on/luhmann2021_ecq_taxonomy.md: appraisal dimensions and the held-out appraisal test.
  • summaries/carry_on/sliwinski2009_stress_bursts.md: stable reactivity against current state.
  • summaries/carry_on/breithaupt2024_retelling_novelty.md: judge-free style markers.
  • summaries/carry_on/sandhan2026_persona_jailbreaking.md: the companion paper read in the same round.
  • docs/research/gss_pilot_design.md: baselines, including the same-answer, same-event other person.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.