Preprint Open access
Pre-training interventions, ex post facto: Grafting model beliefs across checkpoints
Pre-training interventions are critical to alignment research, since beliefs formed during pre-training shape how a model generalizes from later training. One recently popular technique for such interventions is synthetic document fine-tuning (SDF), which aims to alter what the model believes. Ideally, synthetic docume …