Efstathios Stamatatos, Krzysztof Kredens, Piotr Pezik, Annina Heini, Janek Bevendorff, Benno Stein, Martin Potthast. In: Working Notes of CLEF 2023 (Thessaloniki, 18–21 September 2023), CEUR Workshop Proceedings Vol-3497, paper 199; CC BY 4.0. University of the Aegean, Aston University, Bauhaus-Universität Weimar, Leipzig University, ScaDS.AI. Read by researcher R5c (voice), 5 October 2026. Provenance: papers/voice/stamatatos2023_pan_av_overview.provenance.json.
What was read
- Read in full: every line of the pdftotext conversion (1,296 lines, 16 pages): abstract, sections 1–6, Tables 1–7, references. Tables 2, 4 and 5 were re-read in
pdftotext -layoutmode to map rows to numbers. - Not read: the eleven participants' notebook papers (refs 29–39); the PAN 2022 overview (ref 19); the Aston 100 Idiolects Corpus itself.
What they did
- Task: cross-discourse-type authorship verification. Given two texts of different discourse types, give the probability that one person wrote both (0.5 = leave unanswered).
- The first PAN edition with spoken language. Four discourse types from the Aston 100 Idiolects Corpus (English; about 100 people, all native speakers aged 18–22; topics unrestricted): two written (emails, essays) and two spoken (interviews, speech transcriptions). All six pairings of types are used.
- Preparation:
- consecutive emails, and consecutive interview utterances, are concatenated into samples of at least 2,000 characters;
- named entities are replaced by tags, "to reduce the potentially confusing influence of topics";
- non-verbal vocalisations (coughs, laughs) are replaced by tags.
- Split: 56 people for training, a different 56 for testing, similar gender mix. Both sets balanced 50/50 same-author against different-author, also within each pairing.
- Size (Table 2): 8,836 training pairs and 9,656 test pairs. Test pairs by pairing: interview–email 5,214 (54.0%), essay–email 1,618, email–speech transcription 1,074, essay–interview 938, speech transcription–interview 606, essay–speech transcription 206. Average length (test): email 2,346 characters, essay 10,770, interview 2,501, speech transcription 2,537.
- Measures: AUROC; c@1 (accuracy that rewards abstaining); F1; F0.5u (rewards correct same-author answers and abstentions); the complement of the Brier score; the overall score is their mean.
- Baselines: compressor (PPM cross-entropy plus logistic regression); cngdist (cosine over the most common character 4-grams, with two trained thresholds); najafi22 and galicia22 (the two strongest PAN 2022 systems, trained only on written discourse).
- Submissions: 11 teams, 27 runs, executed on TIRA. Most used BERT-family encoders, several with contrastive learning (the approach of Rivera-Soto et al. 2021 is cited); two used classical features (n-grams, function words, vocabulary richness). Only one system treated written and spoken types differently.
Main results (verified)
Overall (Table 4).
- The winner (Ibrahim et al.: Sentence-Transformers with contrastive learning) reached AUROC 0.616, c@1 0.572, F1 0.617, F0.5u 0.562, Brier 0.746; overall 0.623.
- The simple character 4-gram baseline cngdist came fifth of 31, overall 0.595, though with AUROC only 0.516; it answered "same author" 99.8% of the time.
- Every system's AUROC lay between 0.500 and 0.616. The authors: effectiveness is "overall weak resembling in many cases a baseline with random estimates".
By pairing (Table 5, overall score).
- The winner's ES–EM 0.888 and najafi22's ES–EM 0.918 are contaminated: the PAN 2023 essay–email test pairs were part of the PAN 2022 training data, which these systems used.
- No uncontaminated run exceeds 0.64 in any pairing. Best uncontaminated values: ES–EM 0.640, ES–IN 0.605, ES–ST 0.619, EM–ST 0.607, IN–EM 0.613, ST–IN 0.617. Weak runs fall well below 0.5 (down to 0.25).
- Written against spoken is hardest. Averaged over all teams, effectiveness is higher when both texts are written (essay–email) or both spoken (speech transcription–interview) than when one is written and the other spoken. The authors: "the inherent differences between written and spoken language further complicate the task."
- PAN 2022 (written types only: essays, emails, text messages, business memos) was already "extremely difficult"; PAN 2023 is harder.
Other findings.
- Response bias: most strong systems favoured "same author"; only the winner left a moderate share of pairs unanswered (13.9%).
- Runtime ranged from about 4 minutes to 35 hours.
What the corpus examples show (Table 1, same person, two types each).
- The interview is full of fillers and hedges ("erm", "like", "basically", "just", "I don't even know").
- The same person's emails are formulaic ("Many thanks", "Apologies for the late email", "Am I right in assuming").
- The essay is impersonal academic prose; the picture-description transcript is hesitant ("er", "appears to", "I would say").
Limits
- Small population: 56 test authors, one age band (18–22), one country's university students, English only. Pair counts are dominated by interview–email (54%).
- No same-discourse control in the same corpus, so the paper does not measure how much is lost by crossing types. The comparison with earlier fanfiction tasks ("relatively high accuracy") involves other data.
- Samples are short (about 2,000–2,500 characters, about 400–500 words, except essays). Longer samples were not tested.
- The named-entity tagging removes some topical cues but not topic itself.
What it means for Kurisutina (inference)
- A person's "voice" is not one surface style. The same individual writing an email and talking in an interview looks to the best available verifiers almost like two different people (AUROC at most about 0.62 on samples of about 400–500 words). Register — written against spoken, formal against casual, audience — dominates the surface signal.
- The voice test must be register-matched. A replica's chat replies should be compared with the person's own chat or speech, its letters with the person's letters. A pooled "one voice" score would mostly measure the register mix.
- Data collection must cover each register the replica will use, with a few thousand words per register at least. Speech must be transcribed in a way that keeps fillers, hesitations and false starts: they are part of spoken voice, and the corpus kept them (with vocalisations as tags).
- What stays constant across registers is the open question, and the field's best tools barely find it. A slot that must carry voice across registers needs the person's register-switching itself: how they move between formal and casual, which is person-specific (inferred; not tested here).
- Calibration matters for a pass/fail rule. Systems that output hard answers had poor Brier scores even with fair AUROC; a battery should use a calibrated score with an explicit "cannot tell" band, as c@1 and F0.5u reward.