Back to Paper
cs.CRcs.AI
Local ID: 2604.00387v2
AI Summary: gemma4:e4b
RAGShield: Detecting Numerical Claim Manipulation in Government RAG Systems
By KrishnaSaiReddy Patil
Revision History Timeline
v14/1/2026
4/1/2026
“8 pages, 8 tables, 2 figures”
v24/4/2026
4/4/2026
“12 pages, 15 tables, 1 figure, 2 algorithms”
★ Version indexed in ExplorerComparing v1 vs v2
Green = Added • Red = Removed
Title Comparison
RAGShield: Provenance-Verified Defense-in-Depth AgainstDetecting KnowledgeNumerical BaseClaim PoisoningManipulation in Government Retrieval-Augmented GenerationRAG Systems
Authors Comparison
No author changes.
v1 Comment
“8 pages, 8 tables, 2 figures”
v2 Comment
“12 pages, 15 tables, 1 figure, 2 algorithms”
Abstract Word Diff
RAGRetrieval-Augmented Generation (RAG) systems are deployed across federal agencies for citizen-facing services aretax vulnerableguidance, tobenefits knowledgeeligibility, baseand poisoninglegal attacks,information, where adversaries injecta malicioussingle documentsincorrect tonumber manipulatecauses outputs.direct Recentfinancial workharm. demonstratesThis thatpaper asproves fewthat asall 10embedding-based adversarialRAG passagesdefenses canshare achievea 98.2%fundamental retrievalblind successspot: rates.changing Wea observetax thatdeduction RAGby knowledge$50,000 baseproduces poisoningcosine issimilarity structurally0.9998, analogousinvisible to software supply chain attacks, and proposeevery RAGShield,known adetection five-layerthreshold. defense-in-depthAcross framework174 applyingmanipulation supplypairs chainand provenancetwo verificationembedding tomodels, the RAGmean knowledgesensitivity pipeline.gap RAGShieldis introduces:1,459x. (1)The C2PA-inspiredblind cryptographicspot documentis attestationconfirmed blockingon unsignedreal andIRS forgeddocuments.The documentsroot atcause ingestion;is (2)that trust-weightedembeddings retrievalencode prioritizingtopic, provenance-verifiednot sources;numerical (3)precision. aRAGShield formalsidesteps taintthis latticeby withoperating cross-sourceon contradictionextracted detectionvalues catchingdirectly: insidera threatspattern-based evenengine whenidentifies provenancedollar isamounts valid;and (4)percentages provenance-awarein generationgovernment withtext, auditablelinks citations;each andvalue (5)to NISTits SPgoverning 800-53entity compliancethrough mappingtwo-pass acrosscontext 15propagation control(99.8% families.entity Evaluationdetection on a 500-passage Natural Questions corpus with2,742 63real attackIRS documentspassages), and 200 queries against fiveverifies adversaryevery tiersclaim achievesagainst 0.0%a attackcross-source successregistry ratebuilt includingfrom adaptivethe attackscorpus (95%itself. CI:A [0.0%,temporal 1.9%])tracker withflags 0.0%value falsechanges positivethat rate.fall Weoutside honestlyknown reportgovernment thatupdate insiderschedules. in-placeOn replacement430 attacks achievegenerated 17.5%from ASR,real identifyingIRS thedocument fundamentalcontent, limitRAGShield ofdetects ingestion-timeevery defense.one The(0.0% cross-sourceASR, contradiction95% detectorCI catches[0%, subtle1%]) numericalwhile manipulationembedding-based attacksdefenses thatmiss bypass79-90% provenanceof verificationthe entirely.same attacks.