Kurisutina

Learning universal authorship representations

Rafael A. Rivera-Soto, Olivia Elizabeth Miano, Juanita Ordonez, Barry Y. Chen, Aleem Khan, Marcus Bishop, Nicholas Andrews. Proceedings of EMNLP 2021, pages 913–919 (ACL Anthology 2021.emnlp-main.70). Lawrence Livermore National Laboratory and Johns Hopkins University. Read by researcher R5c (voice) for the Kurisutina research batch, 5 October 2026. Provenance: papers/voice/riverasoto2021_luar.provenance.json.

What was read

  • Read in full: every line of the pdftotext conversion (635 lines, 7 pages): abstract, sections 1–5, acknowledgements, references. Tables 1–3 were re-read in pdftotext -layout mode, because the two-column extraction scrambles them; the numbers below come from the layout view.
  • Not read: Figure 1 as an image (its caption was read); the code at github.com/noa/uar.

What they did

  • Question: do authorship representations learned by a neural network in one domain transfer, zero-shot, to other domains, or are they entangled with domain features (topic, genre)?
  • Model (later known as LUAR): a function mapping a collection of documents by one author to a 512-dimensional vector, such that two collections by the same author have higher cosine similarity than collections by different authors.
    • Supervised contrastive loss, temperature 0.01; mini-batches of 256 collections, two per author for 128 authors.
    • Architecture: a pretrained SBERT (BERT plus attention-weighted mean pooling; all parameters updated), applied to excerpts of 32 consecutive subword tokens from each of w documents; self-attention across the w vectors, max-pooling, a linear projection to 512 dimensions.
    • Training samples w from 1 to 16 (w = ⌈1 + 15x⌉, x ~ Beta(3,1)). At test time w = 16 for Amazon and Reddit; for fanfiction, each story is split into paragraphs and all are used.
  • Data (three anonymous domains).
    • Amazon reviews: the 135K authors with at least 100 reviews; 100K train, 35K evaluation (half of each author's reviews as queries, half as targets).
    • Fanfiction (from the PAN 2020 training set): 278,169 stories by 41,000 authors; evaluation uses the 16,456 authors with exactly two stories (one query, one target).
    • Reddit: 1M authors with at least 100 comments for training; evaluation set from Andrews and Bishop (2019), disjoint from and later in time than the training data.
    • Where timestamps exist, evaluation data are future to training data and all targets future to all queries, "to assess robustness to ephemeral aspects of writing, such as topic".
  • Metrics: ranking of all targets by cosine similarity to each query. R@8 = probability the single correct target is in the top 8; MRR = mean reciprocal rank.
  • Baselines: TF-IDF over words (treated as a proxy for a topic model) and the same pipeline with a convolutional encoder (Andrews and Bishop 2019).

Main results (verified)

Table 1, zero-shot transfer (R@8 / MRR; rows = evaluation domain, columns = training domain), proposed model P:

Evaluated on Trained on Reddit Trained on Amazon Trained on fanfiction
Reddit 65.61 / 50.37 23.72 / 15.43 12.10 / 7.64
Amazon 68.91 / 55.59 82.54 / 68.99 28.91 / 20.06
Fanfiction 41.58 / 30.61 36.35 / 26.45 50.89 / 41.20
  • TF-IDF in-domain: Reddit 10.34 / 6.77, Amazon 31.61 / 24.86, fanfiction 31.37 / 22.53. Convolutional in-domain: Reddit 56.32 / 42.38, Amazon 74.30 / 60.60, fanfiction 47.98 / 39.02. The convolutional row for Reddit-evaluated, Amazon-trained prints R@8 6.30 with MRR 9.70 (as printed).
  • Transfer is very uneven. The Reddit-trained model reaches more than 80% of each domain-specific model's R@8 (68.91 of 82.54 on Amazon; 41.58 of 50.89 on fanfiction). Amazon- and fanfiction-trained models transfer poorly (to Reddit: 23.72 and 12.10).
  • Why Reddit transfers. Partly size: a Reddit model trained on only 100K authors reaches 36.7 R@8 on fanfiction, about the Amazon model's 36.35. Adding Amazon authors with fewer reviews (250K authors with 75 reviews; 500K with 50) gained only about 1% each. The authors' explanation: topic diversity. Reddit authors write about many topics, so the model must rely less on topic; the gap of the neural models over TF-IDF is largest on Reddit.
  • Multi-domain training (Table 2): adding Reddit always helps; adding fanfiction always hurts transfer. Best Amazon result 84.84 (Amazon ∪ Reddit, +2.30 over in-domain); best fanfiction result 57.51 (fanfiction ∪ Reddit, +6.62).
  • Random subword masking (Table 3, Amazon as source): masking each training subword with probability 0.45 raised Reddit-target R@8 from 23.72 to 26.36 (MRR 15.43 to 17.50); in-domain and fanfiction changed by about ±1. "Random masking by itself is not sufficient to overcome domain-specific biases." Masking whole words gave similar results. Masking only topic words by frequency (after Stamatatos 2018) was tried but thresholds were hard to set.
  • Conclusion (authors): neural authorship models "do not in general capture universal authorship features"; training data must be controlled to learn the desired invariances. They propose explicitly disentangling content and style as future work.

Limits

  • Retrieval among many authors (tens of thousands of candidates) is the task; there is no calibrated same/different-author decision and no threshold analysis.
  • No experiment on how much text per author is needed: query and target sizes are fixed (16 short documents, or a whole story). No experiment on time gaps beyond requiring targets to be later than queries.
  • Domains differ in genre and topic at once; the paper cannot say which one breaks transfer. Its topic-diversity explanation is an inference from TF-IDF gaps, not a controlled test.
  • Fanfiction pairs are two stories by one author, possibly in different fandoms; no human reference is given.
  • One model family; English only; anonymous internet text, not speech.

What it means for Kurisutina (inference)

  • A learned authorship embedding is a usable ruler for "same author", but not a content-free one. Trained on one genre it partly learns the genre and the topics. The Reddit-trained model is the best general ruler of the three precisely because its authors switch topics. A voice battery that uses such an embedding must (a) choose a model trained on topically diverse data and (b) still control topic in its comparisons, because the ruler is not shown to ignore topic.
  • Scale of signal: with 16 short comments per side, the correct author is in the top 8 among tens of thousands about two-thirds of the time in-domain (Reddit 65.61) and about 40–70% across domains. Sixteen short texts already carry a strong individual signal; a person's voice is not a subtle effect at this sample size.
  • The temporal split is the design to copy. Queries earlier, targets later: the voice test should compare the replica's text with the person's held-out, later or separately elicited texts, never with the texts used to build the slot.
  • For the synthetic world M1, an embedding trained on real people cannot read a 4,096-token synthetic language; M1 needs its own style measure (rates of the generated features, or an embedding trained on the synthetic training persons).

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.