Stefan Evert, Thomas Proisl (FAU Erlangen-Nürnberg), Thorsten Vitt, Christof Schöch, Fotis Jannidis, Steffen Pielström (Universität Würzburg). Proceedings of the NAACL-HLT Fourth Workshop on Computational Linguistics for Literature, pages 79–88, Denver, 4 June 2015 (ACL Anthology W15-0709). Read by researcher R5c (voice), 5 October 2026. Provenance: papers/voice/evert2015_delta.provenance.json.
Why this paper and not Burrows (2002). Burrows's original Delta paper (Literary and Linguistic Computing 17: 267–287) has no open copy; the publisher's site sits behind a bot check, which the project's rule forbids getting around. The same group's 2017 journal version (Digital Scholarship in the Humanities, open access) sits behind the same check. This open conference paper defines Burrows's Delta formally, replicates the main comparison of Delta variants, and explains why they work. Nothing below is taken from Burrows (2002) directly.
What was read
- Read in full: every line of the pdftotext conversion (2,544 lines, 10 pages). Most lines are plotted points from Figures 1–8 extracted as glyphs; the prose, formulas, Tables 1–2, figure captions, footnotes and references were all read.
- Not read: Figures 1–8 as images (curves are described from the prose and captions).
What Delta is (verified, as defined here)
- Each text D is represented by the relative frequencies f_i(D) of the n_w most frequent words (mfw) of the whole collection.
- Each frequency is standardised across the collection: z_i(D) = (f_i(D) − μ_i)/σ_i. This is Burrows's intent "to treat all of these words as markers of potentially equal power".
- Burrows's Delta Δ_B is the Manhattan distance between two z-profiles, Σ|z_i(D) − z_i(D′)|.
- Quadratic Delta Δ_Q is the squared Euclidean distance (Argamon 2008).
- Cosine Delta Δ∠ is the angle between the two z-profiles (Smith and Aldridge 2011).
- For unit-length vectors, Euclidean distance is a monotone function of the angle, so Δ_Q and Δ∠ differ only in normalisation, not in metric.
What they did
- Corpora: three collections, English (1838–1921, Project Gutenberg), French (1827–1934) and German (19th to early 20th century, TextGrid). Each has 25 authors with 3 novels each (75 texts).
- Task: cluster the 75 texts into 25 groups (partitioning around medoids). Quality is measured by the chance-adjusted Rand index (ARI). The number of features n_w varied from 10 to 10,000.
- Experiments:
- Scaling: z-scores against raw relative frequencies.
- Normalisation: the profile vectors scaled to unit length (L1 or L2) before taking distances.
- Supervised feature selection: recursive feature elimination with a linear support-vector classifier, checked on two held-out sets of novels.
Main results (verified)
- Cosine Delta is best and most robust. Its ARI is above 90% "for a wide range of n_w", so most novels group correctly by author, and it degrades slowly up to 10,000 features.
- Burrows's and Quadratic Delta are equal up to about 500 mfw. Quadratic Delta then deteriorates.
- For Burrows's and Quadratic Delta, quality "is substantially diminished for n_w > 5,000" in all three languages.
- The best n_w "depends on many factors – language, text type, length of the texts, quality and preprocessing … and cannot be known a priori". About 2,000 mfw is "a robust and nearly optimal number" for Cosine Delta.
- Standardisation is necessary. Without z-scores, words beyond mfw rank 100 contribute almost nothing, and clustering is poor. Alternative scalings, including Argamon's probabilistic one, did worse.
- Burrows's Delta is robust partly by accident. Its Manhattan distance gives less weight to noisy lower-frequency words, and strongly down-weights words concentrated in few texts (character names, sub-genre vocabulary). In practice it partly protects against content words.
- Vector normalisation is "the key factor".
- Normalising the profiles makes Burrows's Delta as good as Cosine Delta. It hardly matters which norm is used.
- The authors' explanation: an author's style is the pattern of positive and negative deviations from the collection's norm. Different texts by one author express that pattern with different strength (vector length); normalisation removes the strength and keeps the pattern.
- Supervised selection found small feature sets (English 246, French 381, German 234 words) that classify perfectly in cross-validation (SVC 0.99–1.00) and cluster at ARI 0.966–1.000 (Table 1).
- On 71 unseen novels by 19 of the German authors, SVC and MaxEnt reached 0.97 accuracy.
- On 155 unseen novels by 34 authors (6 seen), the 234 selected features gave SVC 0.84, MaxEnt 0.90 and Cosine Delta ARI 0.871. The full feature set gave MaxEnt 0.95 and ARI 0.835 with 2,000 mfw (Table 2).
- The authors read this as features that generalise "fairly well" but are "author-dependent to some extent".
- Selected features are not only function words. They include content words, Roman numerals (chapter counts) and historical spellings, which are corpus artefacts.
Limits
- Whole novels (tens of thousands of words each), three per author, closed sets of 25 authors: an easy setting for stylometry. Nothing here says how Delta behaves on short texts, on speech, or across genres.
- Clustering, not verification: there is no threshold for "same author", and no calibration.
- Only word unigrams; the explanation of why normalisation helps is offered as a conjecture.
What it means for Kurisutina (inference)
- Delta's logic fits the base and slot design. A person's voice, on this account, is their signed deviation profile from a population norm, over the commonest words: more "the", fewer "I", more "but". That is exactly "the slot holds the divergences from the population" (P1, base specification D). The reference norm must be the right population: the culture and register the person belongs to, not a generic corpus.
- Cosine Delta on standardised function-word frequencies is a cheap, transparent, judge-free voice distance. Its main tuning risk (n_w) is tamed by normalisation; about 2,000 mfw is a defensible default for long samples. It needs no neural model, so it also works in the synthetic world M1, where pretrained style embeddings cannot read the synthetic language.
- But it was shown on novels, not on people's daily text. For a person's chat or speech, samples are short, so fewer and more frequent words must be used, and the person's own retest distribution must be measured rather than assumed.
- Content leakage happens even here. Supervised selection picked content words and corpus artefacts. For a voice test, restrict features to frequent function words and punctuation, fixed in advance and never selected on the test data.