Rafael A. Rivera-Soto, Olivia Elizabeth Miano, Juanita Ordonez, Barry Y. Chen, Aleem Khan, Marcus Bishop, Nicholas Andrews. Proceedings of EMNLP 2021, pages 913–919 (ACL Anthology 2021.emnlp-main.70). Lawrence Livermore National Laboratory and Johns Hopkins University. Read by researcher R5c (voice) for the Kurisutina research batch, 5 October 2026. Provenance: papers/voice/riverasoto2021_luar.provenance.json.
What was read
- Read in full: every line of the pdftotext conversion (635 lines, 7 pages): abstract, sections 1–5, acknowledgements, references. Tables 1–3 were re-read in
pdftotext -layoutmode, because the two-column extraction scrambles them; the numbers below come from the layout view. - Not read: Figure 1 as an image (its caption was read); the code at github.com/noa/uar.
What they did
- Question: do authorship representations learned by a neural network in one domain transfer, zero-shot, to other domains, or are they entangled with domain features (topic, genre)?
- Model (later known as LUAR): a function mapping a collection of documents by one author to a 512-dimensional vector, such that two collections by the same author have higher cosine similarity than collections by different authors.
- Supervised contrastive loss, temperature 0.01; mini-batches of 256 collections, two per author for 128 authors.
- Architecture: a pretrained SBERT (BERT plus attention-weighted mean pooling; all parameters updated), applied to excerpts of 32 consecutive subword tokens from each of w documents; self-attention across the w vectors, max-pooling, a linear projection to 512 dimensions.
- Training samples w from 1 to 16 (w = ⌈1 + 15x⌉, x ~ Beta(3,1)). At test time w = 16 for Amazon and Reddit; for fanfiction, each story is split into paragraphs and all are used.
- Data (three anonymous domains).
- Amazon reviews: the 135K authors with at least 100 reviews; 100K train, 35K evaluation (half of each author's reviews as queries, half as targets).
- Fanfiction (from the PAN 2020 training set): 278,169 stories by 41,000 authors; evaluation uses the 16,456 authors with exactly two stories (one query, one target).
- Reddit: 1M authors with at least 100 comments for training; evaluation set from Andrews and Bishop (2019), disjoint from and later in time than the training data.
- Where timestamps exist, evaluation data are future to training data and all targets future to all queries, "to assess robustness to ephemeral aspects of writing, such as topic".
- Metrics: ranking of all targets by cosine similarity to each query. R@8 = probability the single correct target is in the top 8; MRR = mean reciprocal rank.
- Baselines: TF-IDF over words (treated as a proxy for a topic model) and the same pipeline with a convolutional encoder (Andrews and Bishop 2019).
Main results (verified)
Table 1, zero-shot transfer (R@8 / MRR; rows = evaluation domain, columns = training domain), proposed model P:
| Evaluated on | Trained on Reddit | Trained on Amazon | Trained on fanfiction |
|---|---|---|---|
| 65.61 / 50.37 | 23.72 / 15.43 | 12.10 / 7.64 | |
| Amazon | 68.91 / 55.59 | 82.54 / 68.99 | 28.91 / 20.06 |
| Fanfiction | 41.58 / 30.61 | 36.35 / 26.45 | 50.89 / 41.20 |
- TF-IDF in-domain: Reddit 10.34 / 6.77, Amazon 31.61 / 24.86, fanfiction 31.37 / 22.53. Convolutional in-domain: Reddit 56.32 / 42.38, Amazon 74.30 / 60.60, fanfiction 47.98 / 39.02. The convolutional row for Reddit-evaluated, Amazon-trained prints R@8 6.30 with MRR 9.70 (as printed).
- Transfer is very uneven. The Reddit-trained model reaches more than 80% of each domain-specific model's R@8 (68.91 of 82.54 on Amazon; 41.58 of 50.89 on fanfiction). Amazon- and fanfiction-trained models transfer poorly (to Reddit: 23.72 and 12.10).
- Why Reddit transfers. Partly size: a Reddit model trained on only 100K authors reaches 36.7 R@8 on fanfiction, about the Amazon model's 36.35. Adding Amazon authors with fewer reviews (250K authors with 75 reviews; 500K with 50) gained only about 1% each. The authors' explanation: topic diversity. Reddit authors write about many topics, so the model must rely less on topic; the gap of the neural models over TF-IDF is largest on Reddit.
- Multi-domain training (Table 2): adding Reddit always helps; adding fanfiction always hurts transfer. Best Amazon result 84.84 (Amazon ∪ Reddit, +2.30 over in-domain); best fanfiction result 57.51 (fanfiction ∪ Reddit, +6.62).
- Random subword masking (Table 3, Amazon as source): masking each training subword with probability 0.45 raised Reddit-target R@8 from 23.72 to 26.36 (MRR 15.43 to 17.50); in-domain and fanfiction changed by about ±1. "Random masking by itself is not sufficient to overcome domain-specific biases." Masking whole words gave similar results. Masking only topic words by frequency (after Stamatatos 2018) was tried but thresholds were hard to set.
- Conclusion (authors): neural authorship models "do not in general capture universal authorship features"; training data must be controlled to learn the desired invariances. They propose explicitly disentangling content and style as future work.
Limits
- Retrieval among many authors (tens of thousands of candidates) is the task; there is no calibrated same/different-author decision and no threshold analysis.
- No experiment on how much text per author is needed: query and target sizes are fixed (16 short documents, or a whole story). No experiment on time gaps beyond requiring targets to be later than queries.
- Domains differ in genre and topic at once; the paper cannot say which one breaks transfer. Its topic-diversity explanation is an inference from TF-IDF gaps, not a controlled test.
- Fanfiction pairs are two stories by one author, possibly in different fandoms; no human reference is given.
- One model family; English only; anonymous internet text, not speech.
What it means for Kurisutina (inference)
- A learned authorship embedding is a usable ruler for "same author", but not a content-free one. Trained on one genre it partly learns the genre and the topics. The Reddit-trained model is the best general ruler of the three precisely because its authors switch topics. A voice battery that uses such an embedding must (a) choose a model trained on topically diverse data and (b) still control topic in its comparisons, because the ruler is not shown to ignore topic.
- Scale of signal: with 16 short comments per side, the correct author is in the top 8 among tens of thousands about two-thirds of the time in-domain (Reddit 65.61) and about 40–70% across domains. Sixteen short texts already carry a strong individual signal; a person's voice is not a subtle effect at this sample size.
- The temporal split is the design to copy. Queries earlier, targets later: the voice test should compare the replica's text with the person's held-out, later or separately elicited texts, never with the texts used to build the slot.
- For the synthetic world M1, an embedding trained on real people cannot read a 4,096-token synthetic language; M1 needs its own style measure (rates of the generated features, or an embedding trained on the synthetic training persons).