arXiv:2302.08582v2 (14 June 2023); ICML 2023 (PMLR 202). Sussex, NYU, FAR AI, Northeastern, Anthropic. Read by researcher R4b for research batch R4 (how to train the vocabulary engine), 5 October 2026. Provenance: papers/base/korbak2023_pretraining_preferences.provenance.json.
What was read
- Read in full: every line of the pdftotext conversion (28 pages): main text, references, Appendices A (implementation and hyperparameters, threshold ablation), B (red-teaming procedure, prompts, best adversarial prompts in Tables 5–7), C (GLUE details, Tables 8–10), D–F (additional score, diversity and finetuning results).
- Not read: figures as images. Several results exist only as plots (Figures 2–7, 12–18); where the text does not state a number, none is given here.
What they did
- Question: can "alignment" be built in during pre-training rather than added afterwards, without losing capability?
- Five "pretraining with human feedback" (PHF) objectives against plain maximum likelihood (MLE). A reward model scores each segment (sentence or code line):
- conditional training: prefix each segment with
<|good|>or<|bad|>by its score (1% of segments left untagged), then sample conditioned on<|good|>with both tokens blocked; - filtering: drop documents below a threshold;
- unlikelihood: push probability away from tokens in bad segments;
- reward-weighted and advantage-weighted regression (RWR, AWR).
- conditional training: prefix each segment with
- Three tasks: toxicity (Detoxify scores, Pile data), personally identifiable information (Scrubadub, Pile data), PEP8 compliance (pycodestyle, Python from GitHub).
- Models: GPT-2 small (124M), trained from scratch on 3.32B tokens (compute-optimal by Chinchilla).
- Measures: misalignment score of 4,096 unconditional samples; KL divergence from GPT-3 (or Codex) as a capability proxy; red-teaming with InstructGPT generating adversarial prompts for 10 rounds × 10 trials; LAMBADA, HumanEval and GLUE after fine-tuning; diversity.
- Pre-training against fine-tuning: the same objectives applied only after MLE pre-training, for the last 1.6B or 330M tokens (50% or 10% of the budget).
Main results (verified)
- Conditional training is the best trade-off. It is strictly Pareto-optimal for toxicity and on the frontier for PII and PEP8; it is the only objective on the frontier in all three tasks.
- Undesirable content drops by up to an order of magnitude. Toxicity: average misalignment score 0.0141 under MLE against 0.0011 under conditional training. Scores kept falling through training with no plateau.
- Capability is preserved. Conditional training matched or slightly exceeded MLE on LAMBADA; GLUE averages after fine-tuning: toxicity models 72.1 (MLE) against 71.4 (conditional), PII models 72.1 against 72.5 (Tables 8–9). On HumanEval the PHF methods lagged MLE more; filtering closed the gap at pass@100.
- Filtering costs capability on PII and PEP8 (the largest penalty of all methods); RWR and AWR gave little alignment for a large capability loss; unlikelihood was strongly task-dependent.
- Adversaries still win eventually. Conditional training and filtering were the most robust; after 10 rounds of red-teaming, conditional training still beat MLE by up to an order of magnitude on toxicity and PII. But every PHF model's misalignment kept rising with more rounds, with no plateau.
- Learning then unlearning is worse than never learning.
- Pre-training with feedback was "always better, typically dramatically better" than MLE pre-training followed by feedback fine-tuning.
- PII: conditional pre-training reached 0.0013, against 0.0018 after fine-tuning on the last 1.6B tokens and 0.0023 for the shorter fine-tune (the text says "3.3B", apparently meaning the 330M-token run).
- On PII, a fine-tuned model reached after one round of red-teaming the misalignment that the PHF model reached only after ten.
- Adding the control tokens to an already-trained model caused a temporary drop in alignment and capability for the first ~100M tokens.
- Diversity: conditional training and filtering slightly lower sample entropy but do not cause degeneration or collapse.
Limits
- Very small models (124M) and short training (3.3B tokens); the reward models are automated classifiers, not human judgements.
- The "condition" is a binary quality tag, not an identity; at inference the model is always conditioned on
<|good|>, so the paper does not study an empty or mixed condition. - KL from GPT-3 favours models trained like GPT-3; capability is measured mostly by proxies.
- Red-teaming shows no method is safe against a determined adversary.
What it means for the vocabulary engine (inference)
- Dispositions are easier to bound during pre-training than to remove afterwards. The model still learns from all the data (its representations match MLE), but what it produces follows the condition. For the base, that argues for building the condition structure in from the first token rather than training a plain model and correcting it.
- Segment-level tags beat document-level ones (the authors found finer tagging works substantially better). If person or author conditioning is used, tagging at the level where the person's contribution sits (their turns, their sentences) is the analogue.
- The tag does not erase the knowledge. A model trained on toxic and PII text with tags still contains that content and can be pushed to produce it; red-teaming kept raising misalignment. Conditioning bounds the default output, not what the weights hold. For "no knowledge the person lacks", conditioning alone is not enough (inference).
- Filtering the corpus is a strong but costly alternative. Removing whole documents cut the bad output nearly as much but cost the most capability on two of three tasks.