Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:
ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Back to Paper
cs.CLcs.CR

Local ID: 2604.26506v2

AI Summary: gemma4:e4b

SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts

By Yuan Xin, Yixuan Weng, Minjun Zhu, Ying Ling, Chengwei Qin, Michael Backes, Yue Zhang, Linyi Yang

Revision History Timeline

v14/29/2026
4/29/2026

“10 pages, 3 figures, 9 tables”

v25/28/2026
5/28/2026

“17 pages, 5 figures, 8 tables”

★ Version indexed in Explorer

Comparing v1 vs v2

Green = Added • Red = Removed

Title Comparison

SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts

Authors Comparison

Removed:Michael Hahn
Unchanged:Yuan Xin, Yixuan Weng, Minjun Zhu, Ying Ling, Chengwei Qin, Michael Backes, Yue Zhang, Linyi Yang

v1 Comment

“10 pages, 3 figures, 9 tables”

v2 Comment

“17 pages, 5 figures, 8 tables”

Abstract Word Diff

As Large Language Models (LLMs) are increasingly integrated into academic peer review, their vulnerability to adversarial promptshidden --prompts, i.e., adversarial instructions embedded in submissions to manipulate outcomes -- emergesoutcomes, asposes a critical threat to scholarly integrity. To counter this, weWe propose SafeReview, a novelco-evolutionary adversarial training framework wherefor defending LLM-based peer review systems against such attacks. SafeReview jointly trains a Generator model, trainedmodel to create sophisticated attack prompts, is jointly optimizedprompts withand a Defender model taskedto withpreserve theirreview detection.integrity Thisunder systemadversarial manipulation. The Generator is trainedoptimized usingto aproduce lossincreasingly functioneffective inspiredprompt byinjections, Informationwhile Retrievalthe GenerativeDefender Adversarialis Networks,strengthened whichthrough fosterspreference-based atraining dynamicto co-evolutionmaintain betweenconsistent thereviews twobetween models,clean forcingand theattacked Defendersubmissions. toExperimental developresults robustshow capabilitiesthat againstSafeReview continuouslyimproves improvingrobustness attackagainst strategies.adaptive Theprompt resultinginjection frameworkattacks, demonstratesbetter significantlypreserves enhancedpaper resilienceranking tounder novelattack, and evolvinggeneralizes threatsacross attacker architectures compared towith static defenses,defenses. therebyThese establishingresults ademonstrate criticalthe potential of co-evolutionary training as a foundation for securing the integrity ofLLM-assisted peer review.
View Full Version History on arXiv