Back to Paper
cs.LGcs.CLcs.CR
Local ID: 2604.27019v3
AI Summary: gemma4:e4b
Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry
By Wenhao Lan, Shan Li, Xinhua Lai, Meiqi Wu, Junbin Yang, Haihua Shen, Yijun Yang
Revision History Timeline
v14/29/2026
4/29/2026
No submitter comment provided.
v25/17/2026
5/17/2026
No submitter comment provided.
v35/25/2026
5/25/2026
No submitter comment provided.
★ Version indexed in ExplorerComparing v2 vs v3
Green = Added • Red = Removed
Title Comparison
Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry
Authors Comparison
Added:Xinhua LaiMeiqi Wu
Unchanged:Wenhao Lan, Shan Li, Junbin Yang, Haihua Shen, Yijun Yang
v2 Comment
No comment for this version.
v3 Comment
No comment for this version.
Abstract Word Diff
Safety-aligned language models must refuse harmful requests without collapsing into broad over-refusal, yetbut it remains unclear how dynamic adversarial fine-tuning changes therefusal-control internalcarriers: carriersKullback--Leibler of(KL)-constrained refusal.directions or small subspaces that causally modulate refusal without large safe-prompt distribution shifts. We study onea 7B backbone under supervised fine-tuning (SFT) and under Robust Refusal Dynamic Defense (R2D2), a HarmBench-style adversarial fine-tuning procedure that repeatedly refreshes harmful training cases with current jailbreak attacks. Our protocol aligns fixed-sourcealigning HarmBench, StrongREJECT, and XSTest withevaluations awith five-anchor refusal-geometrygeometry suite,measurements, causal interventions, and a sparse adaptive stress test.tests. R2D2 drives fixed-source HarmBench attack success to zero at early checkpoints,checkpoints; buthowever, thatthese regimecheckpoints coincidesalso withexhibit maximal XSTest refusal and complete failure onfail a benign-utility audit. Later checkpoints partially recover benignutility-facing utilitybehavior while partially reopening attack success. Sparse adaptive attacks sharpen the same frontier: step~50 remains closed undersuccess, bothwith adaptive GCG and AutoDAN, whereas adaptiveattack GCGsuccess ASRrate risesrising to 0.415 at step~250step 250 and 0.613 at step~500.step Geometrically,500. Internally, R2D2 preserves a late-layer admissible refusal-control carrier through step~100step 100 and then relocates the best admissible carrier to an early layer by step~250;layer; SFT relocates earlier whileyet remainingremains less robust. Effective rank remainsstays near 1.24, and SFT exhibitsshows larger principal-angle driftdrift, despitearguing worseagainst robustness.both Causaldimensional interventionsexpansion showand thatdrift late-stagemagnitude R2D2as behaviorsufficient isexplanations. controlledCausal byinterventions support a low-dimensional but utility-coupled carrier. These results support a geometry-reorganization account of R2D2 along a robustness--utility frontier.frontier, without establishing adaptive robustness.