Back to Paper
cs.CRcs.IRcs.LG
Local ID: 2605.19847v2
AI Summary: gemma4:e4b
Auditing Privacy in Multi-Tenant RAG under Account Collusion
By Florian A. D. Burnat
Revision History Timeline
v15/19/2026
5/19/2026
No submitter comment provided.
v25/31/2026
5/31/2026
No submitter comment provided.
★ Version indexed in ExplorerComparing v1 vs v2
Green = Added • Red = Removed
Title Comparison
Auditing Privacy in Multi-Tenant RAG under Account Collusion
Authors Comparison
Removed:Brittany I. Davidson
Unchanged:Florian A. D. Burnat
v1 Comment
No comment for this version.
v2 Comment
No comment for this version.
Abstract Word Diff
Multi-tenant retrieval-augmented generation (RAG)RAG services advertiseoften per-accounttreat differentialthe privacyaccount as the operative leakageprivacy boundary: each account's queries are guaranteed toaccount satisfyreceives $(\varepsilon_{\text{acc}},an δ_{\text{acc}})$-DP$(\varepsilon_{\text{acc}},δ_{\text{acc}})$-DP withretrieval respectguarantee toagainst the tenant index. We identify same-index multi-account collusion as a privacy-boundary failure:show forthat $k$this same-tenantframing accountsunderstates coordinatingleakage againstunder thesame-index tenant'saccount indexcollusion. --For theGaussian operativenoise-then-select regimeretrieval, --$k$ knowncoordinated DPsame-tenant compositionaccounts theorycompose impliesto joint leakage degrades unconditionally at$Θ(\sqrt{k}\,\varepsilon_{\text{acc}})$, ratenot $Θ(\sqrt{k}$\varepsilon_{\text{acc}}$; \cdotwe \varepsilon_{\text{acc}})$give fora Gaussian-noisedmatching retrieval.membership-inference Cross-tenantattack and external collusion matchvalidate the rate only under explicit access-control failure (M4); without M4predicted these$\sqrt{k}$ regimesAUC havetrend zeroin leakagescalar, bytop-$K$, designtrained-embedder, and reduce to an architectural audit, not aproduction-scale DPHNSW audit.settings. We exhibit an attack realizing the rate andthen derivegive a RAG-specific MIA prediction we test empirically. To make this per-account/joint gap auditable, we design the firstverifier-runnable audit protocol that operates against unmodifiedattests RAGnoise-then-select deploymentsretrieval and issues a quantitative $(\textsf{PASS}, \varepsilon_{\text{audit}})$reports verdict$(\textsf{PASS},\varepsilon_{\text{audit}})$ for the retrieval-score channel -- the noise-then-select step thecoalitions per-accountup DPto guaranteea actuallydeclared coverscap --$k_{\max}$, without index disclosure,disclosing pipelinethe redesign,index or model-weight exposure. Generation-channel privacy (LLM output conditioned on selected documents) is a separate audit predicate that should compose with ours; wechanging explicitlythe scoperetrieval itdecision out.rule. The protocol composes generic cryptographic primitives (Merkle ledgers, ZK function-application proofs, Gaussian noise attestations) with six RAG-specific primitives (embedder commitment, index-content vector commitment,claim per-accountis queryretrieval-channel ledger,only: noise-then-selectgeneration-channel attestation,leakage cross-tenantand containmentadversarially proof,robust coalition-size estimator) and supportsestimation bothare closed-formcomplementary audit bounds and Rényi-DP moments-accountant tracking.predicates.