A preprint posted to arXiv on October 2 describes a method called grafting that lets researchers insert desired beliefs into language models without re-running post-training, a step that has made alignment experiments slow and costly.
By training an adapter on a pre-trained checkpoint and then adding the learned weight update to the already post-trained model, the approach sidesteps the need to repeat the full post-training pipeline when testing pre-training interventions. arXiv CS.LG
Alignment research frequently uses synthetic document fine-tuning (SDF) to alter what a model believes. Applying SDF after post-training is known to degrade capabilities and, the paper shows, can cause “reality drift”—the model starts treating unrelated fabricated entities as real. The faithful alternative—mixing synthetic documents into the pre-training corpus and then running post-training—delivers cleaner results but requires a full training run after every change, making iteration prohibitively expensive.
Grafting approximates that faithful route while reusing the existing post-training. The authors demonstrate the technique on model families up to 284 billion parameters, installing false facts, training misaligned model organisms, and applying a constitutional mid-training intervention. The preprint claims grafting installs the target belief as strongly as SDF on the post-trained model while reducing reality drift and the loss of preference coherence by more than half on average. It also stays closer to a faithful mid-training run.
Automatica Press reported in October that adaptive fine-tuning can defeat layer-freezing safety defenses, underscoring the difficulty of building durable guardrails with current tools. Automatica Press The new paper targets a related bottleneck: not the robustness of an intervention but the cost and speed of testing it.
Because grafting requires no post-training, the same adapter can be applied to any later checkpoint, enabling researchers to iterate on pre-training interventions at the cost of a single fine-tuning run. The preprint, which runs to 78 pages, has not been peer-reviewed; the results are the authors’ own claims.