Lotta Kiefer, Brisca Balthes, Christoph Leiter, Yamen Ajjour, Elena Schmidt, Steffen Eger. University of Technology Nuremberg (UTN). arXiv:2608.17979v1 [cs.CL], 18 August 2026 (preprint, not yet peer reviewed). Read by researcher R5c (voice), 5 October 2026. Provenance: papers/voice/kiefer2026_style_drift.provenance.json.
What was read
- Read in full: every line of the pdftotext conversion (1,883 lines, 16 pages): abstract, sections 1–6, limitations, ethics, acknowledgements, references, Appendices A–I (example texts with translations, feature configuration, training setup, full results, standardised results, AI-era results, most and least stable features, pairwise feature analysis, licences). Figure 2's numeric cells and Table 3 were re-checked in
pdftotext -layoutmode. - Not read: Figures 1, 3, 4 and 8 as images (captions and legends read; Figure 4's per-year values are not in the text, only its end points as stated in the prose); the released data and code.
What they did
- Question: authorship verification (AV) assumes an author's style is stable enough to tell them from others. How much do changes of genre, time, and the arrival of generative-AI writing tools break that?
- Data (AVShift, German): scraped from one platform, fanfiktion.de, which hosts three genres per user: forum posts ("General Chit Chat"), reviews of others' stories, and fanfiction chapters. Aligned texts by the same authors in two or three genres, 2004–2025 (21 years). Texts under 50 words dropped, over 3,000 words truncated; URLs and sign-offs removed; topical vocabulary kept, with topic controlled in pair construction instead (positive pairs from different works, different reviewed stories, different threads). At most 5 positive and 5 negative pairs per author; authors split 80/10/10 with no overlap.
- GenreShift: three in-domain sets (forum, review, story), three cross-genre sets (one text from each genre), and a mixed set. 15,272–40,320 pairs per set; 1,849–4,806 users; mean sample length 173 words (forum), 184 (review), 1,358 (story).
- TimeShift: ten slices by the gap between the two texts' dates, 0–12 months up to 108–120 months. Positive and negative pairs share the gap. Downsampled per genre to forum 670, review 446, story 4,466 pairs per slice.
- AIShift: four eras (2004–2010, 2011–2016, 2017–2022, 2023–2025); leave one era out (4,122 training and 1,374 test pairs per split).
- Verifiers (three paradigms):
- XGBoost on more than 4,000 handcrafted features (most frequent word bigrams, part-of-speech trigrams, character 4-grams, emojis, word-length distribution, punctuation, message length, out-of-vocabulary rate), using the element-wise difference of the two vectors;
- MSR, a fixed multilingual style embedding (Kim et al. 2025; 36 languages, 13 domains), with only the threshold tuned (Youden's J);
- Gemma-4-31B-it fine-tuned with LoRA to answer "same author?" (four H200 GPUs).
- Metric: macro F1 on balanced pairs (accuracy in the appendix); paired bootstrap with 10,000 resamples. A standardised variant fixes every sample at 500 words (concatenating or truncating) and equalises training size (3,640 samples). An English check on CrossNews (news articles and tweets by the same authors).
- Feature stability score: 1 − σ_within(f) / σ_between(f), comparing a feature's variation within an author across genres with its variation between authors.
Main results (verified)
Time (TimeShift).
- F1 falls steadily as the gap between the two texts grows, in every genre.
- Correlations of F1 with the gap: review r = −0.94 (ρ = −0.95), story r = −0.88 (ρ = −0.93), forum r = −0.75 (ρ = −0.76); all p < 0.05.
- Review: F1 0.90 for texts 0–12 months apart, 0.69 for texts 9–10 years apart (−0.21, the paper's "up to 0.21 F1").
- Forum: 0.76 to 0.68.
- "A large decline occurs already after the first year." The ranking of genres (review above story above forum) holds at every gap.
- The authors conclude that "verification is only reliable when comparing documents written within relatively short time windows without further model modifications."
Genre (GenreShift, Figure 2, F1).
- Best in-domain: review 0.89, story 0.80, forum 0.78. Reviews carry the strongest authorial signal and forum posts the weakest. The ranking survives standardisation to 500 words and equal training size: review 0.84, story 0.74, forum 0.74.
- Cross-genre, best model (Gemma trained on the mixed set): review–forum 0.77, story–forum 0.70, review–story 0.72. The figure's "Average" row (over its 21 model rows): 0.63, 0.47 and 0.53.
- A single-genre model loses about 0.2 F1 out of genre. On reviews, Gemma scored 0.89 trained on reviews, 0.71 trained on stories and 0.68 trained on forum posts.
- Diverse training data restores much of the loss. The mixed-trained Gemma beat the specialised cross-genre models in every cross-genre setting.
- Standardised to 500 words, the advantage disappears: cross-genre F1 dropped to 0.67–0.68 for all genre pairs; Gemma no longer led, and MSR or XGBoost matched it. The authors attribute this to the smaller training sets.
- The fixed style embedding (MSR) collapses on some genre pairs: F1 0.07–0.54 on story–forum and 0.14–0.65 on review–story, depending on the calibration genre. Its threshold, set in one genre, does not transfer.
What survives genre change (handcrafted features).
- In t-SNE, author vectors separate by genre more than by author.
- Feature stability ranged from −1.70 to 0.90 (mean 0.22, median 0.25); 80% of features were positive.
- By pair (Table 3): review–forum mean 0.31 (85% positive), story–forum 0.19 (75%), review–story −0.06 (48%).
- The prose attributes the −0.06 and 48% to story–forum. The table assigns them to review–story, which is also the pair ranked second in performance, as the prose says. This summary follows the table.
- The most stable features are function-word bigrams and part-of-speech patterns. Examples: "aber in", "so als", "sollte man", "ist doch", "nicht gerade", "zum Beispiel", plus punctuation habits ("!!!!", "|", "#", "+").
- The least stable are mostly character 4-grams tied to narrative past tense and genre vocabulary ("hatt", "sah", "Blic", "ckte"), plus average message length.
- Different genre transitions destabilise different features: overlap of the 100 most stable features between pairs is 5–27%, of the 100 least stable 2–43%. Stability rankings still correlate (Spearman 0.42–0.72).
AI era (AIShift). There is no consistent chronological pattern; AI-era texts were not harder to verify (review reached its best F1, 0.89, on the AI era). This is not a test of AI use, which was not annotated.
English check (CrossNews). Same pattern. Gemma reached F1 0.90 on tweet–tweet and 0.87 on article–tweet (+0.11 and +0.07 over the earlier best); MSR reached 0.88 on article–article.
Limits
- One German hobby platform; mostly young fan-community writers; three written genres, no speech. The forum, review and story genres differ in length as well as style (the standardised runs address length).
- The time analysis uses one model per genre (the best in-domain model) and small slices (446–670 pairs for review and forum). Figure 4's intermediate values are not reported in text.
- "Time" mixes ageing of the writer, changing platform norms and changing topics. The design controls era by giving negatives the same gap, but it does not separate an individual's drift from shared language change.
- Preprint; a few text–table inconsistencies (above).
What it means for Kurisutina (inference)
- A person's written style is a moving target over years, detectably so within one year. A verifier that recognises the person at 0.90 F1 within a year recognises them at about 0.69 after a decade. For a replica:
- the voice reference must be dated, and the slot should be trained mainly on recent text, or weighted by date (as the Roman bot's late-period bias showed);
- a test that compares the replica with texts from many years back will under-rate a faithful replica.
- "The person's own consistency" is a function of time gap and genre, not one number. The human reference for a voice test must be computed at the same gap and in the same genre as the replica's comparison: the person against themselves at gap Δ, in genre g.
- What carries across genres is the habitual small stuff: function-word pairings, part-of-speech habits and punctuation quirks. Narrative or genre vocabulary does not. These are the features a slot must reproduce for voice to survive a change of context, and the features a judge-free measure should weight.
- Fixed embeddings need per-genre calibration. MSR's threshold set in one genre failed in another. A battery that uses a pretrained style embedding must calibrate its decision rule per register, on the person's own and other people's texts in that register.
- Diverse data per person helps verification, and probably the slot too. Training a verifier on varied genres made it robust; by analogy (untested), a slot trained on the person's text in several registers should learn what is constant in their voice better than one trained on a single register.