Tuhin Chakrabarty (Stony Brook University), Jane C. Ginsburg (Columbia Law School), Paramveer Dhillon (University of Michigan; MIT Initiative on the Digital Economy). arXiv:2510.13939v4 [cs.CL], 17 March 2026 (first version October 2025; preprint; preregistered at OSF, osf.io/zt4ad). Read by researcher R5c (voice), 5 October 2026. Provenance: papers/voice/chakrabarty2025_author_finetuning.provenance.json.
What was read
- Read in full: every line of the pdftotext conversion (3,100 lines, 50 pages): main text, methods, data and code statements, ethics, references, and the supplement S1–S10 (author list, writing prompt, fine-tuning pipeline, evaluation interface, AI detection, cliché density, model specifications, cell counts, coefficient tables 8–19, per-author results, corpus-size regression, mediation, costs, deviations from preregistration).
- Not read: Figures 1–25 as images (captions read; their key numbers are restated in the text). Examples of emulations and reader rationales (Figures 15–23) exist only as images. The GitHub data and the OSF preregistration were not opened.
What they did
- Task: write an excerpt of up to 450 words "emulating the style and voice" of one of 50 acclaimed authors (Nobel, Booker and Pulitzer winners and emerging writers; some read in translation, using one translator per author). Each excerpt follows a detailed content specification taken from a real excerpt by that author.
- Human writers: 28 MFA students or graduates from top US programs, 3 per author. Their prompt had 20 sample excerpts, a written description of the author's style, and the content specification. No time limit; $75 per excerpt.
- AI condition 1, in-context: GPT-4o, Claude 3.5 Sonnet and Gemini 1.5 Pro received the identical prompt (150 human–AI pairs).
- AI condition 2, fine-tuned (30 living authors):
- Data: each author's complete works (purchased ePubs; 0.89M tokens for Tulathimutte, 2 books, up to 10.9M for Pamuk), cut into 250–650-word excerpts.
- Training pairs ("instruction back-translation"): GPT-4o wrote a content description of each excerpt. GPT-4o was fine-tuned through the API on pairs of "write an n-word excerpt about [content] emulating the style and voice of [author]" → the real excerpt.
- Settings: 1 epoch for four prolific authors, 3 for the rest. The target excerpts were excluded from training.
- Post-processing: outputs were resampled if they reproduced training text verbatim (ROUGE-L 0.16–0.23 against the originals), then lightly corrected by GPT-4o for grammar and typos.
- 90 human–AI pairs.
- Readers: 28 MFA-trained readers (never judging their own work) and 516 college-educated general readers (Prolific; US and UK; at least 30 seconds per judgment; AI-written rationales excluded).
- Blind pairwise forced choice on two outcomes. Writing quality: two excerpts shown. Stylistic fidelity: the two excerpts plus the original author excerpt as reference.
- Each choice needed a 2–3-sentence rationale.
- 10,920 judgments in all.
- Analysis: logistic models with CR2 cluster-robust errors by reader, and Holm correction. The preregistration specified mixed models, which did not converge; the deviation is reported.
- AI detectors: Pangram and GPTZero, threshold 0.9. Stylometric mediators: cliché density (LLM-proposed, intersected over two authors' annotations), readability, adjectives.
Main results (verified)
In-context imitation loses to human experts; fine-tuned imitation beats them (Table 9, odds ratios, AI against human).
| Readers | Fidelity, in-context | Quality, in-context | Fidelity, fine-tuned | Quality, fine-tuned |
|---|---|---|---|---|
| MFA-trained | 0.16 [0.08, 0.29] | 0.13 [0.06, 0.28] | 8.16 [4.69, 14.2] | 1.87 [1.16, 3.02] |
| College-educated | 1.06 [0.90, 1.25] | 1.82 [1.53, 2.15] | 16.65 [13.03, 21.28] | 5.42 [4.41, 6.66] |
- Predicted probability that experts choose the AI excerpt for fidelity: in-context GPT-4o 0.32, Claude 0.38, Gemini 0.17; fine-tuned GPT-4o 0.74. General readers: 0.45–0.54 in-context, 0.80 fine-tuned.
- Judges matter.
- Experts agree moderately: κ = 0.58 for fidelity and 0.41 for quality in-context; 0.54 and 0.67 fine-tuned.
- General readers barely agree: κ = 0.10 and 0.08 in-context; 0.26 and 0.12 fine-tuned.
- General readers do not see stylistic fidelity: in-context they were at parity on fidelity (OR 1.06) while experts strongly preferred the humans. In-context they also preferred AI for quality.
- Quality and fidelity judgments correlate at r = 0.64 for experts in-context, falling to 0.39 after fine-tuning; for general readers, 0.52 to 0.45 (SI S2.2, Figure 14).
- Per author (Table 13; 72 trials each, about 87.5% general readers).
- Fidelity: 29 of 30 fine-tuned authors above parity; median win rate 0.83 (IQR 0.70–0.94); 26 significant. Range 0.98 (Roxane Gay) to 0.18 (Tony Tulathimutte).
- Tulathimutte has the smallest corpus (0.89M tokens). The authors read his case as "certain idiosyncratic voices resist algorithmic mimicry".
- Quality: 28 of 30 above parity.
- The "fine-tuning premium" over in-context ranged from −12.2 to +67.4 points for fidelity (median +34.2) and from −14.1 to +46.1 for quality (median +13.4).
- No corpus-size effect in this range (0.89M–10.9M tokens). Slope on fidelity 0.0089 per million tokens (CI −0.019 to 0.037), R² = 0.015; Spearman ρ = −0.040. The premium also does not correlate with size (r < 0.1).
- Detectors catch prompted imitation, not fine-tuned imitation.
- Pangram flagged 97% of in-context excerpts and 3% of fine-tuned ones; GPTZero 91% and 0%. No human excerpt was flagged.
- Among experts in-context, each unit of detector score cut the odds of choosing the excerpt by a factor of about 6.3 (β = −1.85). Fine-tuning removed most of this penalty (interaction β = +2.56 for style).
- Cliché density correlated with the detector score at r = 0.60. It mediated 16.4% of the detector effect on expert preference in-context and 1.3% (not significant) after fine-tuning; all three features together 25.4% against −3.2%.
- The authors' reading: fine-tuning "reduces rather than merely masks" the model's own stylistic signature, the "officious disembodied robovoice". A fine-tuned model reproduced Tulathimutte's "ribald and somewhat profane" register and Díaz's Spanish–English slang.
- Cost: fine-tuning $22.25–$272.50 per author (median $77.88), plus $3 to generate 100,000 words.
- Minor inconsistency: the author list marks 29 fine-tuned authors, but Rachel Cusk also appears among the 30 in the per-author results.
Limits
- The comparison is imitation against imitation. MFA writers imitating an author compete with a model imitating the author. Nobody tested whether readers could tell the model from the author themselves, which is the replica question.
- Famous authors with 0.9–11 million tokens of polished published prose, almost certainly known to the base model. The training prompt names the author. Literary excerpts only; no conversation, speech or everyday writing.
- Fine-tuned outputs were resampled and grammar-corrected by GPT-4o before evaluation, which helps quality and may erase idiosyncrasies.
- One fine-tuned model (GPT-4o, through OpenAI's API; adapter type not disclosed). Readers judge one 450-word excerpt at a time.
- Partly a legal-economic paper: about a third of the main text concerns fair use.
What it means for Kurisutina (inference)
- The clearest evidence for P2 in the voice domain. The same frontier models, given 20 real excerpts and an expert description of the voice in context, imitated poorly: experts preferred the human imitator 6–8 to 1, and detectors flagged 97%. Trained on the author's text, imitation beat experts and evaded detection. Voice carried by retrieved examples is weak and reads as the base model's voice; voice trained into weights replaces the base model's voice. For the replica, the person's voice must be in the slot's weights; reading the person's episodes from the store will not supply it and should not be relied on to.
- The base model's own voice is the drift signal. Cliché density and AI-detector scores tracked the in-context outputs' failure. A replica whose language drifts toward the base's habits (clichés, polish, upbeat safety) is losing the person. AI-text detectors and cliché density are cheap drift alarms, not proof of fidelity: fine-tuned text passed the detectors and could still be wrong for the person.
- Instruction back-translation is a usable training recipe for a voice slot. Pair a neutral description of what was said with the person's actual wording of it, so the adapter learns how the person says things given the content. This separates voice from content by construction. For a real person, do not "fix" their grammar or typos: those are voice.
- Human judging of voice needs the right judges. Experts agreed moderately (κ about 0.5–0.67); lay readers barely agreed (κ about 0.1–0.26), could not see fidelity, and rewarded fluency. For a person, the experts are the people who know their voice: partner, family, close friends. A human voice test should use close witnesses with the person's own text as the reference, and should report agreement among judges.
- The battery must test against the person, not against other imitators. The right comparison is replica against the person's own held-out text, with judges asked "which one is the person?" plus a person-against-person control (two real samples), so the human reference is the rate at which judges can tell the person from themselves.