Kurisutina

Same author or just same topic? Towards content-independent style representations

Anna Wegmann, Marijn Schraagen, Dong Nguyen (Utrecht University). Proceedings of the 7th Workshop on Representation Learning for NLP (RepL4NLP 2022), pages 249–268, Dublin, 26 May 2022 (ACL Anthology 2022.repl4nlp-1.26). Read by researcher R5c (voice), 5 October 2026. Provenance: papers/voice/wegmann2022_content_independent_style.provenance.json.

What was read

  • Read in full: every line of the pdftotext conversion (2,992 lines, 20 pages): abstract, sections 1–7, ethics, acknowledgements, references, Appendices A–F (hyperparameters, development results, STEL details, error analysis, clustering, compute, intended use). All tables (1–13) were re-read in pdftotext -layout mode, since reading order scrambles them; numbers below come from the layout view.
  • Not read: Figures 1–2 as images (captions and the example texts embedded in them were read); the code and data on GitHub (nlpsoc/Style-Embeddings).

What they did

  • Problem: style embeddings are increasingly trained on authorship verification (AV: same author or not?). But an author tends to write about the same topics, so a good AV score may come from content, not style.
  • Two changes to the training task.
    1. Contrastive AV (CAV): given an anchor A1, choose which of A2 (same author) and B (different author) shares A1's author; trained with a triplet loss. Plain AV pairs are trained with a contrastive loss.
    2. Content control (CC): the different-author utterance B is drawn (i) from the same conversation as A1, (ii) from the same subreddit ("domain"), or (iii) at random ("no" control).
  • Data: a 2018 Reddit sample, 100 subreddits, 600 conversations each with at least 10 posts; authors split 70/15/15 into train, development and test with no overlap. Per CC level: 210k training triplets (420k AV pairs), 45k development and 45k test triplets. About 195k–270k authors per training set; one author is A1 at most 9 times.
    • Same-author pairs (A1, A2) shared a conversation 27% and a subreddit 56% of the time. Different-author pairs shared a subreddit 1–2% of the time under no control, and always shared a conversation under conversation control (Table 1). Without control, topic is therefore a strong shortcut.
  • Models: siamese BERT-uncased, BERT-cased and RoBERTa-base (Sentence-Transformers), 4 epochs, batch 8, margin 0.5 (margins 0.4–0.6 barely mattered), three seeds. Cosine similarity throughout. About 846 GPU hours in all.
  • Evaluation. (a) AV (AUC) and CAV (accuracy) on test sets at each CC level. (b) STEL: 1,830 tasks on four known style dimensions (formal/informal 815, complex/simple 815, number substitution "nb3r" 100, contractions 100), built from paraphrase pairs so content cannot help. (c) A new STEL-Or-Content task: the anchor must be matched either to a sentence with the same style but different content or to its own paraphrase in the other style; accuracy is the share of choices that follow style. (d) Agglomerative clustering of 14,756 test utterances.

Main results (verified)

Table 2 (test set, RoBERTa, mean ± SD over seeds). AV AUC on conversation / domain / random different-author pairs:

  • untrained RoBERTa-base: .53 / .57 / .61;

  • AV-trained, no control: .58 / .63 / .79;

  • AV-trained, domain control: .68 / .71 / .73;

  • AV-trained, conversation control: .69 / .70 / .71;

  • CAV-trained, conversation control: .69 / .70 / .71 (CAV accuracy .68 / .69 / .70).

  • The uncontrolled model's apparent skill is largely topic. It is the best on random pairs (.79) and the worst on same-conversation pairs (.58): a drop of 0.21 AUC when the impostor is on the same topic. The conversation-controlled model is nearly flat across test sets (.69–.71).

  • Same-conversation pairs are the hardest test for every model, as expected if content cues are removed.

  • Single utterances verify weakly. On development data, the best AV accuracy at the AUC-optimal threshold was 0.64 on conversation-controlled pairs (RoBERTa, conversation-trained), 0.65 on subreddit-controlled pairs and 0.71 on random pairs (RoBERTa, uncontrolled) (Appendix Table 7).

Table 3 (STEL and STEL-Or-Content accuracy, RoBERTa).

Model STEL all STEL-Or-Content all Formal o-c Complex o-c nb3r o-c Contraction o-c
RoBERTa-base, untrained .80 .05 .09 .01 .13 .00
AV, conversation control .71 .35 .64 .13 .04 .00
AV, no control .72 .22 .46 .03 .05 .00
CAV, conversation control .71 .42 .69 .24 .03 .04
CAV, no control .71 .24 .50 .04 .06 .00
  • Untrained language-model embeddings follow content almost always: RoBERTa-base chose the same-content sentence over the same-style one in 95% of STEL-Or-Content tasks (accuracy .05).
  • Content control helps, but content still wins more often than style. The best model (CAV with conversation control) chose by style in 42% of tasks, below the 0.5 random line; the authors note the task is very hard because the false choice shares most words with the anchor.
  • Fine-tuning on AV lowered plain STEL (.80 to about .71): formality and contraction were kept or improved, complexity and number substitution dropped. Error analysis found many STEL complex/simple items ambiguous (50 of 55 "unlearned" items, 29 of 41 "learned" ones).
  • What the embedding groups (clustering, 7 clusters): 46.2% of same-author pairs fell in the same cluster against 20.1% expected by chance. Visible consistencies were surface habits: no final punctuation (97% of one cluster against 37% elsewhere), lower-case "i" and missing apostrophes ("didnt", "thats"), the typographic apostrophe ’ against ' (90% against about 0%), and line breaks (72% against 22%). Untrained RoBERTa's clusters separated mainly length and URLs.

Limits (including the authors')

  • Conversation labels exist only in conversational data; even within one conversation, content can still identify an author ("my husband" against "my wife").
  • STEL covers four dimensions only, two of them tiny (100 items); STEL-Or-Content carries the "triplet problem" (the paraphrase may sometimes be stylistically closer).
  • Training on thousands of authors may teach group styles (age, education, community), not only individual ones.
  • Single Reddit utterances; English; no time dimension; no spoken language; no human reference for how often a person's own utterances "match".

What it means for Kurisutina (inference)

  • Content leakage is the central threat to a voice test, and it is large. An embedding trained without topic control lost 0.21 AUC when the impostor was on the same topic. A replica that talks about the person's topics, with the person's names and facts from its store, would score as "same author" on such an embedding without having the person's voice. The battery must use impostors on the same topic (the same prompts answered by other people), not random other texts.
  • Use a content-controlled style embedding, and still control content in the design. Conversation-controlled training is the best available recipe in this paper, but even it follows content more often than style when content is held equal by paraphrase. No embedding alone can certify "voice, not content"; the test design must hold content fixed (same prompts, same topics for person, impostors and replica).
  • The habits these models pick up are exactly the low-level, unconscious markers a slot must carry: punctuation and capitalisation habits, apostrophe type, contractions, line breaking, utterance length. They are cheap to count directly and can be reported beside any embedding score.
  • Aggregate, don't judge single utterances. One short utterance verifies at about 0.64 accuracy against a same-topic impostor (0.71 against a random one); a voice verdict needs many utterances per sample (LUAR used 16 per side).

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.