A non-peer-reviewed preprint posted October 2 reports that a layer-freezing safety defense failed to preserve refusal when models were fine-tuned on 100 harmful examples across six checkpoints from four model families.

The study challenges the assumption that pinpointing the layers responsible for a model's safety behavior can yield a durable guardrail against fine-tuning attacks. Earlier research had shown that a few dozen harmful examples can strip refusal and had localized safety-related behavior to specific layers, directions, and tokens.

The preprint, by Jungseob Lee, Dongyub Jude Lee and colleagues, tests whether that localization survives an adaptive attacker. The authors first replicated a recovery method that patches clean hidden states into a compromised model, which restored refusal at a reproducible transition depth. When they froze every layer up to that depth and repeated the attack, refusal remained near zero on all six checkpoints, with recovery transitions moving above the frozen boundary.

A second study removed the top two singular directions from the fine-tuning update, restoring refusal after short attention-only fine-tunes on four checkpoints. On Llama-3.1-8B, however, ordinary training changes weakened the repair, and an attacker who spread the update defeated it. A spectral detector calibrated on benign Llama fine-tunes missed most repair failures on that checkpoint.

The authors write that localized freezing can still help when a few harmful examples enter training data unintentionally. But the results, they add, show an adaptive attacker can bypass a region identified by recovery and defeat a repair that works across multiple checkpoints, motivating five checks for defenses against adaptive fine-tuning. Code is available at the link in the paper. The preprint has not been peer-reviewed and does not include independent replication results.