Abstract
Pre-training interventions are critical to alignment research, since beliefs formed during pre-training shape how a model generalizes from later training. One recently popular technique for such interventions is synthetic document fine-tuning (SDF), which aims to alter what the model believes. Ideally, synthetic documents would be mixed into pre- or mid-training, but every change to a pre-training corpus must be followed by a full post-training run before its effect can be measured, making iteration slow and expensive. Common practice instead applies SDF to an already post-trained model. This is known to leave artifacts and degrade capabilities, and, as we show, it makes the model treat fabricated entities unrelated to the documents as real, a failure we call reality drift. We propose grafting: train the SDF adapter on the pre-trained checkpoint, then add the learned weight update to the post-trained model, which approximates the faithful approach while reusing the existing post-training. We demonstrate this by installing false facts, training misaligned model organisms and applying a constitutional mid-training intervention, across model families up to 284B parameters. Grafting installs the target belief as strongly as SDF on the post-trained model while reducing both reality drift and the loss of preference coherence by more than half on average, and it stays closer to a faithful mid-training run. Because grafting requires no post-training, the same adapter can be applied to any later checkpoint, enabling researchers to iterate quickly on pre-training interventions at the cost of a single fine-tuning run.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Nutter, P., Roytburg, D., Dumas, C., Ou, J., & Feng, S. (2026). Pre-training interventions, ex post facto: Grafting model beliefs across checkpoints. https://omanscience.com/en/articles/pre-training-interventions-ex-post-facto-grafting-model-beliefs-across-checkpoints
MLA 9
Nutter, Peter, et al. "Pre-training interventions, ex post facto: Grafting model beliefs across checkpoints." https://omanscience.com/en/articles/pre-training-interventions-ex-post-facto-grafting-model-beliefs-across-checkpoints.
Chicago (author–date)
Nutter, Peter, Dani Roytburg, Clément Dumas, Jinghua Ou, and Shi Feng. 2026. "Pre-training interventions, ex post facto: Grafting model beliefs across checkpoints." https://omanscience.com/en/articles/pre-training-interventions-ex-post-facto-grafting-model-beliefs-across-checkpoints.
Harvard
Nutter, P., Roytburg, D., Dumas, C., Ou, J. and Feng, S. (2026) 'Pre-training interventions, ex post facto: Grafting model beliefs across checkpoints', Available at: https://omanscience.com/en/articles/pre-training-interventions-ex-post-facto-grafting-model-beliefs-across-checkpoints.
Vancouver
Nutter P, Roytburg D, Dumas C, Ou J, Feng S. Pre-training interventions, ex post facto: Grafting model beliefs across checkpoints. https://omanscience.com/en/articles/pre-training-interventions-ex-post-facto-grafting-model-beliefs-across-checkpoints
IEEE
P. Nutter, D. Roytburg, C. Dumas, J. Ou, and S. Feng, "Pre-training interventions, ex post facto: Grafting model beliefs across checkpoints," https://omanscience.com/en/articles/pre-training-interventions-ex-post-facto-grafting-model-beliefs-across-checkpoints.