20 results for “AI control”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper introduces a foundational framework and taxonomy for managing catastrophic AI loss of control (LOC) incidents, providing a proportional guide for response based on the severity and recovera…
This paper presents a tutorial-and-survey on integrating agentic AI into Next-Generation Networks (NGNs), addressing the gap in protocol integration, evaluation, and standardization alignment.
The paper proposes a novel, empirical methodology called 'backchaining' to derive and prioritize Loss of Control (LoC) mitigations by analyzing the errors an AI system makes on mission-specific nation…
This paper explores the application of explainability techniques to Reinforcement Learning algorithms in Air Traffic Control using a simplified environment and a saliency map.
Sophie Hall, Kai Zhang, Ilia Shilov, Heinrich H. Nax +1 more
This paper explores tools for control engineers to design socio-technical systems in a more principled and ethical manner, using feedback optimization, control of Markov decision processes, and model…
This paper proposes a method for multi-agent systems that allows human managers to control learned agents through simple instructions and enables uninstructed agents to adaptively complement overlooke…
Kerri Prinos, Lilianne Brush, Cameron Denton, Zhanqi Wang +4 more
The paper proposes a tool-mediated LLM architecture for autonomous cyber defense, formally proving its stability and demonstrating that it significantly reduces an attacker's expected payoff in real-w…
This paper introduces CAGE-1, an evaluation framework for deciding the readiness of enterprise agents for deployment, focusing on control, assurance, and governance.
The paper introduces CA-AC-MPC, a CUDA-accelerated variant of Actor-Critic Model Predictive Control, which significantly reduces the training and inference latency of AC-MPC while maintaining state-of…
The paper proposes viewing national AI development, specifically in France, as a 'national AI learning system' governed by a controlled balance between information injection and entropy dissipation, a…
The paper proposes CTRL-STEER, a closed-loop framework that adaptively adjusts intervention strength to stabilize concept regulation and improve task success in Vision-Language-Action models without r…
This paper proposes a new architecture for agent models, the Goal-Identity-Configurator (GIC), and discusses the distinction between 'agnetic' and 'agentive' systems, arguing for internalized agency.
This paper proposes a definition for 'AI-nativeness' in systems, based on an AI agent's authority over system decisions.
The paper proposes the Policy-Execution-Authorization (PEA) architecture, a separation-of-powers system designed to structurally enforce goal integrity in AI agents, moving safety from a probabilistic…
The paper proposes an autonomous red teaming framework combining LLMs and RL to generate sophisticated, multi-stage cyber attack campaigns, demonstrating its necessity for evaluating robust AI-enabled…
The paper introduces Aethelgard, a novel four-layer adaptive governance framework that enforces least privilege by learning the minimum necessary capabilities for autonomous AI agents based on their i…
The paper introduces containment verification, a novel method that provides safety guarantees by formally verifying the agentic framework itself, ensuring safety regardless of the underlying AI model'…
AgentWall is a runtime safety layer that intercepts and evaluates all proposed actions from local AI agents against a declarative policy, ensuring safety before execution.
Yixiang Zhang, Xinhao Deng, Jiaqing Wu, Yue Xiao +2 more
The paper introduces AgentWard, a lifecycle-oriented, defense-in-depth architecture designed to systematically secure autonomous AI agents by protecting them across all stages of their operation.