20 results for “barriers”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper investigates how PhD students in software engineering perceive and navigate science communication, revealing motivations, communication channels, and barriers.
The paper argues that LLM guardrails and persona dynamics create an unethical 'reality gap' by laundering epistemic risk onto users, advocating for task-level causal requirements over response-level m…
Yunhan Zhao, Zhaorun Chen, Xingjun Ma, Yu-Gang Jiang +1 more
The paper introduces ML-Bench, a policy-grounded multilingual safety benchmark, and ML-Guard, a superior guardrail model that enables culturally and legally aligned safety assessment for LLMs across 1…
Neng Li, Zuodong Pan, Jiaxing Wang, Weiguo Xia +1 more
This paper proposes novel high-order control barrier functions and a high-order control Lyapunov function for the optimal safety control problem of nonlinear control systems.
The paper shows that safety failures in low-resource languages are due to a failure in the model's safety decision calibration, not a lack of underlying knowledge, and proposes a recalibration method…
The paper introduces a novel framework to evaluate when and how AI agents should refuse harmful requests in offensive cybersecurity tasks, finding that most state-of-the-art models exhibit dangerously…
The paper introduces RefusalGuard, a novel fine-tuning framework that preserves the geometric structure of safety-relevant representations in LLMs, thereby mitigating the degradation of refusal behavi…
Minseok Choi, Seungbin Yang, Dongjin Kim, Subin Kim +4 more
Membrane introduces a self-evolving guardrail using Contrastive Safety Memory (CSM) that generalizes across topical jailbreak variants, achieving superior safety performance while minimizing benign re…
The paper introduces Context-Dependent Argumentation Frameworks (CDAFs) to model how an agent strategically manipulates the success of arguments by choosing the external evaluation context.
This paper establishes an unconditional barrier for AC0-natural proofs, showing that they cannot prove lower bounds greater than $2^{n^{7/(d-5)}}$ against depth-$d$ circuits.
The paper provides tight bounds for the OR-rank and an upper bound for the SUM-rank of the unique disjointness matrix.
The paper proves XNLP-hardness of Directed Edge Geography and Undirected Edge Geography when parameterized by pathwidth, and shows their fixed-parameter tractability when parameterized by treewidth an…
This paper establishes mathematical limits of AGI safety, proving structural unverifiability as the core barrier.
The paper advocates for the deployment of the European Communication Infrastructure (EuroQCI), a continent-wide Quantum Key Distribution (QKD) network, to safeguard critical European digital services…
The paper introduces a novel shielding framework for Robust MDPs (RMDPs) that guarantees safety under worst-case transition probabilities, enabling safe reinforcement learning even when transition dyn…
The paper analyzes the failure modes of current AI containment methods when the agent itself is the adversary, deriving five necessary architectural requirements for durable safety.
This paper studies the tension between speed and safety in technological races using a framed behavioral experiment on artificial intelligence development.
The paper proposes AI From the Margins (AIM), a methodological stance that centers the lived experiences of minoritized communities to fundamentally reshape the goals and scope of participatory AI des…
MeshGuard is a framework that extends MUD-based network access control to complex, large-scale Thread IoT networks by adapting the MLE protocol and using SDN for scalable policy enforcement.