Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:
ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Back to Paper
cs.CRcs.AI

Local ID: 2604.23238v2

AI Summary: gemma4:e4b

Hiding in Plain Sight: Detectability-Aware Antidistillation of Reasoning Models

By Max Hartman, Vidhata Jayaraman, Moulik Choraria, Yash Savani, Lav R. Varshney

Revision History Timeline

v14/25/2026
4/25/2026

No submitter comment provided.

v25/8/2026
5/8/2026

No submitter comment provided.

★ Version indexed in Explorer

Comparing v1 vs v2

Green = Added • Red = Removed

Title Comparison

Protecting theHiding Trace:in APlain PrincipledSight: Black-BoxDetectability-Aware ApproachAntidistillation Againstof DistillationReasoning AttacksModels

Authors Comparison

Added:Yash Savani
Unchanged:Max Hartman, Vidhata Jayaraman, Moulik Choraria, Lav R. Varshney

v1 Comment

No comment for this version.

v2 Comment

No comment for this version.

Abstract Word Diff

Frontier models push the boundaries of what is learnable at extreme computational costs, yet distillationDistillation via sampling reasoning traces exposes closed-source frontier models to adversarial third parties who can bypass their guardrails and misappropriate their capabilities, raising safety, security, andcapabilities. intellectualAntidistillation privacymethods concerns.aim Toto address this, there is growing interest in building antidistillation methods, which aimthis toby poisonpoisoning reasoning traces to hinder downstream student model learning while maintainingpreserving teacher performance. However, current techniques lack theoreticalmethods grounding,overlook requiringdetectability, eitherboth heavysemantic fine-tuningand orsyntactic, accesswhich toerodes studenttrust modelin proxiesthe forteacher's gradientoutputs basedand attacks,signals andthe oftendefense's leadpresence to a significantadversaries. teacherWe performanceaddress degradation.this Ingap thisby work,formulating weantidistillation presentas a theoretical formulationStackelberg ofgame antidistillationwhose asconstraint aset Stackelbergexplicitly game,encodes groundingdetectability, aand problemshow that hasperturbing sosparingly faroffers largelyan beeneffective, approachedless heuristically.detectable Guidedalternative byto poisoning the desiredfull designtrace. propertiesDrawing ouron formulationmechanistic reveals,interpretability, we proposeidentify \texttt{TraceGuard},thought ananchors, efficient,sentences post-generationwith black-boxdisproportionate methodcounterfactual toinfluence poisonon sentencesmodel withoutputs, highas importancea forprincipled teachersparse reasoning.target: Ourcritical workto offersreasoning ayet scalableminimally solutiondetectable. toWe shareinstantiate modelthis insightsin safely,TraceGuard, ensuringa thattraining-free, theblack-box advancementproof-of-concept ofthat reasoninglocates capabilitiesthought doesanchors notvia comebranching-token atdetection theand costpoisons ofthem intellectualto privacydegrade orstudent AIdistillation safetywhile alignment.preserving trace coherence.
View Full Version History on arXiv