Kurisutina

Customizing large language model generation style using parameter-efficient finetuning (StyleTunedLM)

Xinyue Liu, Harshita Diddee, Daphne Ippolito (Carnegie Mellon University). Proceedings of the 17th International Natural Language Generation Conference (INLG 2024), pages 412–426, 23–27 September 2024 (ACL Anthology 2024.inlg-main.34; also arXiv:2409.04574). Read by researcher R5c (voice), 5 October 2026. Provenance: papers/voice/liu2024_styletunedlm.provenance.json.

What was read

  • Read in full: every line of the pdftotext conversion (1,592 lines, 15 pages): abstract, sections 1–5, limitations, ethics, availability statement, references, Appendices A–D (authors, the GPT-4 prompt, linguistic features, fine-tuning details, qualitative samples, per-author tables 6–9, masking samples). Tables 1–4 were re-checked in pdftotext -layout mode.
  • Not read: Figures 1–5 as images (perplexity bars, cosine-similarity bars, t-SNE, confusion matrices); the prose gives their headline values. The demo site and released checkpoints were not opened.

What they did

  • Question: can a low-rank adapter (LoRA), trained by next-token prediction on one author's unstructured text, make a pretrained model write in that author's style, better than prompting with examples?
  • Setup:
    • One LoRA adapter per author on Llama-2-7B. Learning rate 5×10⁻⁵, 3 epochs, 256-token chunks, two A6000 GPUs. The LoRA rank is not reported.
    • Ten authors from Project Gutenberg: Richardson, Austen, Hawthorne, Twain, Wilde, Gilman, Woolf, Vernon Lee, Wodehouse, Orwell. Each author's books are split by whole book into training, validation and test.
    • The authors say the target users are writers "who already have some 1k–50k tokens of prior work".
  • Baselines: the base model prompted with 5 or 10 random 256-token excerpts of the author ("5shot", "10shot"); Llama-2-7b-chat told the author's name and asked to continue in their style ("instruct").
  • Evaluation: 100 prompts (50 written by GPT-4, 50 sentence openings from the test books); 256-token continuations.
    • Perplexity on the author's held-out text.
    • Style-embedding similarity: a Sentence-Transformer trained (pairwise loss) to separate the ten authors' training excerpts; cosine similarity of generations to each author's mean embedding.
    • A 10-way BERT author classifier (chance 0.10).
    • Linguistic alignment at three levels: lexical (nouns, verbs, adjectives and unique words per sentence, subjectivity, concreteness; MSE); syntactic (distribution of simple, compound, complex and compound-complex sentences; Jensen–Shannon divergence); surface (commas, semicolons, colons, words per sentence, word length; MSE).
  • Extras:
    • Content control: person names masked from the loss during training, using spaCy tags.
    • Data size: 5%, 35%, 70% and 100% of 80k training tokens.
    • Instruction following: merging the style adapter with an adapter trained on the LIMA instruction set.

Main results (verified)

  • The adapter beats prompting with examples (Table 2, averages over the ten authors).

    Method Classifier accuracy Lexical MSE Syntactic JSD Surface MSE
    StyleTunedLM 0.879 1.39 0.06 2.04
    5-shot 0.693 3.80 0.07 5.43
    10-shot 0.680 3.31 0.06 4.68
    instruct (name only) 0.263 2.67 0.15 3.78
    • Ten examples in context were no better than five.
    • Telling a chat model the author's name gave the worst style match.
    • Per author the picture is mixed (Table 6). For Hawthorne, the instruct baseline had lower lexical and surface error than the adapter.
  • Perplexity on the author's held-out text fell by 7.0% on average, and by 13.6% for Richardson, the oldest and most unlike modern English.

  • Data size (Table 3, averages; 100% = 80k tokens, 3 epochs).

    Training text Classifier accuracy Cosine similarity Lexical MSE Surface MSE
    5% (about 4k tokens) 0.11 (chance 0.10) 0.57 6.04 10.02
    35% (about 28k) 0.44 0.74 3.49 5.59
    70% (about 56k) 0.81 0.92 1.44 2.27
    100% (80k) 0.88 0.95 1.39 2.04
    • Perplexity barely moved (13.47 to 12.72).
    • The authors cite Eder (2015) for 5,000–10,000 words as the minimum for stable attribution.
  • Content control by masking names (Table 7, prompts built to elicit names).

    • The share of generated names that also occur in the training books fell: Wodehouse 0.50 to 0.23, Austen 0.61 to 0.45, Gilman 0.46 to 0.12.
    • The authors judge the effect on style "minimal". The 10-way classifier's accuracy fell for every author, however: Austen 1.0 to 0.76, Hawthorne 1.0 to 0.72, Orwell 0.98 to 0.16. Embedding similarity fell less (Orwell 1.0 to 0.76).
  • Merging with an instruction adapter kept style (Woolf:LIMA weight 1:1 gave cosine 0.70, against 0.57 for LIMA alone) and left benchmark scores flat (MMLU 0.336–0.340).

Limits (including the authors')

  • The authors are famous and long dead; their books are almost certainly in Llama-2's pretraining data. The adapter may unlock a style the base already knows. The authors say results "might not extend" to low-resource writers, "such as a short essay".
  • The rulers are trained on the same ten authors' books, so they measure separation among ten known authors, not resemblance to a person against the whole population. Their sensitivity to names suggests they also read content.
  • Literary prose only; no conversation, speech or everyday writing; no human judges; one base model; the LoRA rank is unreported; the data-size runs are averages without intervals.

What it means for Kurisutina (inference)

  • Voice can live in an adapter's weights, and it beats putting examples in the context. The adapter outperformed 5 and 10 in-context examples, and adding examples did not help. This is the P2 pattern for voice: retrieved samples of the person's text in the prompt are a weaker carrier of voice than weights trained on it.
  • Order of magnitude of data: for a style the base has never seen, assume tens of thousands of tokens at least. Here about 4k tokens gave nothing measurable, about 28k gave partial capture, and 56–80k gave most of it, for authors the base had probably already read. For an unknown private person the requirement is probably higher. Collect 100k+ words per person where possible (guess, to be tested).
  • The slot learns content along with style unless told not to. Unmasked adapters reproduced the training texts' characters. Masking names during training is cheap and should be standard for a slot's voice component. P2 adds a reason: names and facts in the voice adapter would carry identity-relevant content into weights by the back door.
  • Evaluate voice with rulers that cannot read content. A classifier that recognised Orwell 98% of the time by his characters' names recognised him 16% of the time without them. The battery's voice measure must be validated against same-topic, different-author impostors before use (Wegmann 2022).
  • Separate adapters for style and capability can be merged. This suggests a slot could keep a voice component separate from other components and combine them by weight ratio, a usable "voice strength" knob (inference from one author, Woolf).

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.