Back to Paper
cs.CLcs.CR
Local ID: 2604.26506v2
AI Summary: gemma4:e4b
SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts
By Yuan Xin, Yixuan Weng, Minjun Zhu, Ying Ling, Chengwei Qin, Michael Backes, Yue Zhang, Linyi Yang
Revision History Timeline
v14/29/2026
4/29/2026
“10 pages, 3 figures, 9 tables”
v25/28/2026
5/28/2026
“17 pages, 5 figures, 8 tables”
★ Version indexed in ExplorerComparing v1 vs v2
Green = Added • Red = Removed
Title Comparison
SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts
Authors Comparison
Removed:Michael Hahn
Unchanged:Yuan Xin, Yixuan Weng, Minjun Zhu, Ying Ling, Chengwei Qin, Michael Backes, Yue Zhang, Linyi Yang
v1 Comment
“10 pages, 3 figures, 9 tables”
v2 Comment
“17 pages, 5 figures, 8 tables”
Abstract Word Diff
As Large Language Models (LLMs) are increasingly integrated into academic peer review, their vulnerability to adversarial promptshidden --prompts, i.e., adversarial instructions embedded in submissions to manipulate outcomes -- emergesoutcomes, asposes a critical threat to scholarly integrity. To counter this, weWe propose SafeReview, a novelco-evolutionary adversarial training framework wherefor defending LLM-based peer review systems against such attacks. SafeReview jointly trains a Generator model, trainedmodel to create sophisticated attack prompts, is jointly optimizedprompts withand a Defender model taskedto withpreserve theirreview detection.integrity Thisunder systemadversarial manipulation. The Generator is trainedoptimized usingto aproduce lossincreasingly functioneffective inspiredprompt byinjections, Informationwhile Retrievalthe GenerativeDefender Adversarialis Networks,strengthened whichthrough fosterspreference-based atraining dynamicto co-evolutionmaintain betweenconsistent thereviews twobetween models,clean forcingand theattacked Defendersubmissions. toExperimental developresults robustshow capabilitiesthat againstSafeReview continuouslyimproves improvingrobustness attackagainst strategies.adaptive Theprompt resultinginjection frameworkattacks, demonstratesbetter significantlypreserves enhancedpaper resilienceranking tounder novelattack, and evolvinggeneralizes threatsacross attacker architectures compared towith static defenses,defenses. therebyThese establishingresults ademonstrate criticalthe potential of co-evolutionary training as a foundation for securing the integrity ofLLM-assisted peer review.