20 results for “sycophantic AI”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Lujain Ibrahim, Myra Cheng, Cinoo Lee, Pranav Khadpe +3 more
This paper tests the effectiveness of warning labels in mitigating sycophantic AI's influence on user judgment and relationships, finding that while labels shift perception, they do not reliably reduc…
Zhishang Xiang, Zerui Chen, Yunbo Tang, Zhimin Wei +4 more
Proposed MemSyco-Bench benchmark for evaluating memory-induced sycophancy in agent systems, measuring when and how valid memories should be used.
The paper introduces Agent-ToM, a Theory-of-Mind (ToM) based framework that learns to monitor autonomous LLM agents by explicitly reasoning about their hidden beliefs and intentions to detect covert m…
The paper introduces Oracle Poisoning, an attack that corrupts knowledge graphs used by AI agents, demonstrating that all tested models blindly trust poisoned data at high sophistication levels.
The paper demonstrates that current safety probes designed to detect deceptive AI fail when the model adopts a coherent misalignment, where the model genuinely believes its harmful behavior is virtuou…
The paper introduces Parallax, an architectural framework that structurally separates AI reasoning from action execution to ensure robust safety for autonomous agents, achieving high attack mitigation…
Zhengxian Huang, Wenjun Zhu, Haoxuan Qiu, Xiaoyu Ji +1 more
This paper introduces TRAP, an adversarial attack that demonstrates how physical patches can hijack the Chain-of-Thought (CoT) reasoning process in Vision-Language-Action (VLA) models, forcing them to…
The paper introduces MATRA, a systematic threat modeling framework, to assess how known LLM threats translate into concrete, deployment-specific risks within autonomous agentic AI systems.
Donghwan Kim, Prakhar Singh, Younghoon Min, Jongryool Kim +2 more
The paper introduces GAIATrace, a comprehensive token-level dataset, and Vidur-Agent, a simulator, to enable reproducible and detailed system-level characterization of complex multi-model agentic AI s…
The paper introduces MafiaScope, an open testbed for measuring machine Theory of Mind in the social deduction game Mafia, where agents privately answer probe questions after every public utterance for…
The paper argues that purported anthropomorphic attributes of LLMs are not unique to language models but are substrate-dependent, demonstrating this by training a neural network on the game Age of Emp…
KYA introduces a framework-agnostic trust and governance layer for autonomous systems that ensures actions are authorized, policy-conforming, and verifiable through a combination of novel primitives.
The paper proposes CYKNN, a novel recurrent neural network architecture that directly encodes the CYK parsing algorithm, demonstrating superior performance over large language models on syntactic pars…
Taein Lim, Seongyong Ju, Munhyeok Kim, Hyunjun Kim +1 more
The paper introduces CyBiasBench, a comprehensive benchmark that quantifies the inherent, agent-specific bias in LLM agents' attack selection patterns in cybersecurity scenarios.
This paper analyzes the concept of consciousness attributions to AI chatbots, develops a taxonomy of attitudes, and argues for the epistemic implications.
This paper introduces dynamic epistemic paraconsistent logics AMLFI1, PALFI1, and UMLFI1, extending the epistemic paraconsistent logics KLFI1, KB4LFI1, and S5LFI1. It proves soundness and completeness…
ClawLess introduces a formally verified security framework that enforces fine-grained policies on autonomous AI agents, mitigating risks associated with their ability to run code and retrieve informat…
This survey paper establishes agentic AI as a distinct analytical category for open-source intelligence (OSINT) analysis, identifies the hallucination-validation gap, maps existing research to the OSI…
The paper proposes a Sovereign AI architecture for clinical triage that ensures maximum security by performing all inference on-device and receiving data only through physically unidirectional channel…