20 results for “safety”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
The paper proposes SafeDIG, a robust safety steering framework that adapts Diffusion Transformers for text-to-image generation by treating safety control as position-aware sparse feature transfer, ens…
Xian Qi Loye, Qinglin Su, Zhexin Zhang, Shiyao Cui +4 more
The paper introduces RUBAS, a rubric-based reinforcement learning framework that improves agent safety by providing fine-grained, multi-dimensional rewards for complex tool-use scenarios.
Hanjiang Hu, Yiyuan Pan, Jiaxing Li, Xusheng Luo +4 more
VLESA is a novel framework that monitors human activities from egocentric video to predict and intervene in dangerous actions by incorporating goal-conditioned safety checks based on inferred intent.
The paper proposes an algorithmic method using conformal prediction to formally certify high-probability safety for Belief-Space Neural Safety Filters (BeliefSF), significantly improving safety guaran…
Yuan Xiong, Linji Hao, Shizhu He, Yequan Wang +1 more
The paper proposes Janus, a framework for long-horizon agent safety using multi-agent simulation and CoAA-RL, which trains guards to anticipate risks and make safety judgments based on both observed p…
Hao Li, Jingkun An, Zijun Song, Pengyu Zhu +7 more
SafeSteer proposes a localized on-policy distillation method that restricts safety alignment to specific safety tokens, thereby achieving strong safety performance with minimal degradation to general…
Yusong Zhao, Yuejin Xie, Youliang Yuan, Junjie Hu +3 more
The paper introduces PaSBench-Video, a comprehensive streaming video benchmark designed to rigorously test multimodal LLMs' ability to issue proactive safety warnings, finding that current models stru…
The paper introduces BOA, a novel framework that measures agent safety by exhaustively searching the entire in-budget trajectory space, thereby identifying unsafe behaviors missed by traditional sampl…
Yunhan Zhao, Zhaorun Chen, Xingjun Ma, Yu-Gang Jiang +1 more
The paper introduces ML-Bench, a policy-grounded multilingual safety benchmark, and ML-Guard, a superior guardrail model that enables culturally and legally aligned safety assessment for LLMs across 1…
This paper evaluates the integration of Language Model Machines (LLMs) into control tasks in automotive contexts from a safety assurance perspective, identifying conceptual and concrete challenges.
This paper proposes a simple real-time monitor for LLMs that turns an external verifier signal into an alarm decision by thresholding, showing competitiveness with advanced monitors in mathematical re…
Xuwei Ding, Skylar Zhai, Linxin Song, Jiate Li +5 more
The paper introduces OS-BLIND, a benchmark demonstrating that current safety evaluations fail to detect critical vulnerabilities in computer-use agents when user instructions are benign, showing high…
Ian Dardik, Yining She, Sam Procter, Keaton Hanna +2 more
This paper introduces FASR, a tool that automates the identification of unsafe control actions (UCAs) in System-Theoretic Process Analysis (STPA) using model-based engineering and robustness analysis.
The paper introduces SafetyDrift, a predictive model that forecasts when AI agents will violate safety protocols by analyzing the cumulative risk across sequences of individually safe actions.
GLiNER Guard (GLiGuard) introduces a unified, efficient encoder family that simultaneously performs safety classification and PII detection in a single forward pass, offering a practical, low-cost alt…
This paper analyzes incidents involving autonomous vehicles (AVs) and identifies gaps in current safety paradigms, suggesting the need for human-interactive autonomy.
The SAFE approach enhances fault-tolerant trust management in VANETs by ensuring vehicles send updated feedback reports before leaving a witness area, significantly reducing erroneous penalization of…
Qi Li, Jiu Li, Pingtao Wei, Jianjun Xu +7 more
This paper comparatively evaluates DKnownAI Guard against three competitors, demonstrating that DKnownAI Guard achieves superior performance in detecting both agent-specific threats and harmful conten…
The paper introduces SafeRedirect, a system-level defense that prevents frontier LLMs from generating harmful content during legitimate tasks that structurally require it, significantly reducing unsaf…