20 results for “Fairness attacks”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
A novel structure-aware reinforcement learning-based method is proposed to exacerbate unfairness in recommender systems by modeling structural and sequential dependencies.
The paper introduces the quotient semivalue mechanism to provide fair data attribution that is resistant to contributors manipulating their reported identities by splitting or duplicating data.
The paper proposes Fair Fine-tuning (FFt), a method that fine-tunes a model using an Equalized Odds constraint on a complementary distribution, and provides a formal theoretical bound linking this fai…
The paper proposes Fair Fine-tuning (FFt), a method that fine-tunes a model using an Equalized Odds constraint on a complementary distribution, and theoretically proves that this approach significantl…
This paper introduces a novel, comprehensive dataset that logs various cheating activities, including difficult-to-detect network flow disruption cheats, for the purpose of developing robust detection…
The paper proposes Test-Time Collective Action (TTCA), a framework allowing groups of users to correct algorithmic biases in black-box systems by applying pooled, proxy-based perturbations at inferenc…
Li Zhang, Yuyuan Li, XiaoHua Feng, Jiaming Zhang +2 more
This paper addresses the challenge of achieving optimal fairness and accuracy simultaneously in multi-class classification by proposing novel in-processing and post-processing algorithms that converge…
This paper is an invited commentary on Ying Cheng's Psychometrika focus article comparing test fairness and algorithmic fairness. The commentary discusses the distinction between equality and equity a…
The paper argues that AI security research is imbalanced, focusing too much on demonstrating attacks and not enough on developing practical, usable defenses.
BiasEdit introduces a training-free framework that automatically detects and edits unknown social biases in web-sourced image datasets to construct a debiased dataset for fair visual classification.
Alexander Nemecek, Osama Zafar, Yuqiao Xu, Wenbiao Li +1 more
The paper argues that current AI content watermarking benchmarks fail to test for bias across different languages, cultures, and demographics, proposing a new set of evaluation standards to ensure fai…
Taein Lim, Seongyong Ju, Munhyeok Kim, Hyunjun Kim +1 more
The paper introduces CyBiasBench, a comprehensive benchmark that quantifies the inherent, agent-specific bias in LLM agents' attack selection patterns in cybersecurity scenarios.
This paper investigates how individual agent biases amplify system-wide unfairness in multi-agent systems, demonstrating that uniform exposure to bias can elevate overall bias beyond the sum of indivi…
Bang An, Yibo Yang, Dandan Guo, Ebtisam Alshehri +2 more
The paper proposes Dual-Reference SFT to mitigate harmful fine-tuning of language models by embedding harmful QA pairs in benign training samples.
The paper analyzes and documents various double-dip reward abuse attacks that exploit flaws in how cashback and reward engines handle transaction refunds, proposing formal invariants and defensive alg…
The paper demonstrates that current defenses against malicious fine-tuning of foundation models are insufficient because they only address fixed attacks, and introduces a unified adaptive attack that…
The paper demonstrates that the current per-token billing model for LLMs is susceptible to systematic overcharging because auditing frameworks must rely on evidence provided by the very companies that…
The paper demonstrates that the current per-token billing model for LLMs is susceptible to systematic inflation because auditing frameworks must rely on evidence provided by the service provider, crea…
Xinpeng Lv, Chunyuan Zheng, Yunxin Mao, Renzhe Xu +8 more
The paper introduces Individual Fairness-aware Strategic Classification (IFSC), a framework that models interdependent strategic manipulation where agents imitate nearby positively decided peers to ac…
Yunrui Yu, Xuxiang Feng, Pengda Qin, Pengyang Wang +4 more
The paper introduces Dummy-Aware Weighted Attack (DAWA), a novel evaluation method that significantly reduces the reported robustness of Dummy Classes-based defenses by simultaneously targeting both t…