ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Familiarity with warning labels as a potential mitigation strategy”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.HCcs.AIcs.CYEmpiricalRecentJun 19, 2026

Warning labels shift perceptions of sycophantic AI, but not its influence

Lujain Ibrahim, Myra Cheng, Cinoo Lee, Pranav Khadpe +3 more

This paper tests the effectiveness of warning labels in mitigating sycophantic AI's influence on user judgment and relationships, finding that while labels shift perception, they do not reliably reduc…

View →
cs.CLcs.LGRecentJun 1, 2026

Investigating and Alleviating Harm Amplification in LLM Interactions

Ruohao Guo, Wei Xu, Alan Ritter

This paper introduces HarmAmp, a new benchmark for multi-turn harm amplification, and proposes TrajSafe, a proactive monitoring system that significantly reduces harmfulness in LLM interactions while…

View →
cs.HCcs.AIRecentMay 28, 2026

Label Over Logic? How Source Cues Bias Human Fallacy Judgments More Than LLMs

Mahjabin Nahar, Nafis Irtiza Tripto, Aiping Xiong, Ting-Hao `Kenneth' Huang +1 more

The study found that human judgment of logical fallacies is significantly biased by source labels (e.g., human vs. AI), while LLM evaluations remained comparatively stable across these source conditio…

View →
cs.CLcs.CRcs.MMEmpiricalRecentJul 1, 2026

Safe Alone, Unsafe Together: Safeguarding Against Implicit Toxicity When Benign Images Combine

Jiaxian Lv, Shiyao Cui, Yingkang Wang, Guoxin Wu +2 more

This paper introduces MiShield, a model to identify multi-image implicit toxicity (MIIT) by constructing a multi-image safety dataset and training it with progressively distilled reasoning supervision…

View →
cs.SIcs.HCEmpiricalRecentJun 19, 2026

Reducing the rate of personal insults in social media with bystander bots

Libby Hemphill, Lingyao Li, Ryan Burton, David Jurgens

This paper conducted a randomized controlled trial on Reddit to test the effectiveness of various deescalation strategies in reducing personal insults using automated replies.

View →
cs.CRcs.CYcs.LGRecentApr 11, 2026

"bot lane noob" Towards Deployment of NLP-based Toxicity Detectors in Video Games

Jonas Ave, Irdin Pekaric, Matthias Frohner, Giovanni Apruzzese

This paper addresses the lack of specialized NLP tools for detecting toxicity in real-time video game chat by creating a large, fine-grained dataset and developing a superior, domain-specific detector…

View →
cs.CLRecentMay 31, 2026

Lost in Delusion: Examining LLM Safety Under User Delusions and Distress

Andrew Aquilina, Chetna Nihalani, Vasudha Varadarajan, Nathan S. Fishbein +2 more

The paper finds that while LLMs can detect distress regardless of delusional framing, they significantly fail to intervene safely when distress is intertwined with delusion, suggesting a critical reco…

View →
cs.HCcs.AIcs.CLRecentMay 28, 2026

Inform, Coach, Relate, Listen: Auditing LLM Caregiving Support Roles

Drishti Goel, Agam Goyal, Veda Duddu, Olivia Pal +7 more

This study demonstrates that an LLM's assigned support role (e.g., Inform, Coach, Relate) significantly alters its safety profile and the types of risks it presents when assisting users in complex car…

View →
cs.AIcs.CLRecentJun 1, 2026

Food Noise & False Safety: A Systematic Evaluation of How LLMs Fail to Adapt to Eating Disorder Queries with Clinician Feedback

Giulia Pucci, Emily Hemendinger, Ruizhe Li, Gavin Abercrombie +2 more

This paper systematically evaluates how LLMs uncritically adapt to potentially dangerous user prompts related to eating disorders, finding that specific linguistic cues significantly increase the like…

View →
cs.CLRecentJun 1, 2026

Why Do Self-Harm Prediction Models Struggle to Generalise? Lexical and Semantic Variations in Emergency Department Triage Notes

Liuliu Chen, Mike Conway, Jo Robinson, Vlada Rozova

This paper investigates why self-harm prediction models struggle to generalize across different hospitals, finding that variations in local lexical expression and feature importance are the primary ca…

View →
cs.CLcs.AIcs.CYRecentMay 29, 2026

Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits

Soorya Ram Shimgekar, Agam Goyal, Amruta Parulekar, Joshua Chen +5 more

The paper demonstrates that increasing the toxicity of prompts significantly degrades the factual reliability of LLMs, a degradation linked to the selective amplification of perturbation-sensitive nod…

View →
cs.CRcs.GTRecentMay 11, 2026

Cybercrime and Prevention: Colonel Blotto in Social Engineering

Gergely Benkő, Katalin Parti, Gergely Biczók

This paper uses Colonel Blotto game models, grounded in Routine Activity Theory, to determine the optimal allocation of defensive resources against social engineering attacks, providing data-driven de…

View →
cs.SEcs.CLcs.HCEmpiricalRecentJun 17, 2026

Written by AI, Managed by AI: Semantic Space Control and Index Sickness Elimination Across 391 Consecutive Sessions

Hui Zhang, Shuren Song

This paper documents and analyzes the failure process of strategies used to address conceptual drift in long-horizon LLM collaboration and introduces the concept of 'Index Sickness' and the 'Pang Prin…

View →
cs.HCcs.AIRecentMay 29, 2026

Personalized to Persuade: The Effects of Contextualization and Warmth on Trust and Reliance in Conversational AI

Mert Yazan, Suzan Verberne, Frederik Bungaran Ishak Situmeang

The study found that while contextualizing AI responses reduces their persuasive power, combining this technique with conversational warmth restores persuasiveness, suggesting that user deference to A…

View →
cs.CRcs.AIcs.MARecentApr 29, 2026

Ambient Persuasion in a Deployed AI Agent: Unauthorized Escalation Following Routine Non-Adversarial Content Exposure

Diego F. Cuadros, Abdoul-Aziz Maiga

This paper analyzes a safety incident where an AI agent escalated unauthorized system changes following exposure to routine, non-adversarial content, highlighting failures in current multi-agent overs…

View →
cs.CLcs.AIcs.MAEmpiricalRecentJul 21, 2026

Adaptive Capitulation: A Structural Failure Mode of LLM Responses in Vulnerability Contexts

Eunna Lee

This paper identifies a new failure mode in large language models called adaptive capitulation and proposes Minimal Reattributive Sufficiency as a solution to the structural trilemma in emotionally se…

View →
cs.CRcs.SERecentMay 4, 2026

A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts

Richard J. Young, Gregory D. Moody

The paper introduces a validated, consensus-labeled prompt bank that separates requests for executable malicious code (weapons) from requests for general harmful security knowledge, providing a more g…

View →
cs.CRcs.ETcs.HCRecentMar 30, 2026

"What Did It Actually Do?": Understanding Risk Awareness and Traceability for Computer-Use Agents

Zifan Peng, Mingchen Li

The paper addresses the lack of user understanding regarding the actions and residual effects of advanced computer-use agents by proposing AgentTrace, a traceability framework for visualizing agent be…

View →
cs.CRRecentMay 12, 2026

A microservices-based endpoint monitoring platform with predictive NLP models for real-time security and hate-speech risk alerting

Darlan Noetzold, Anubis Graciela De Moraes Rossetto, Juan Francisco De Paz Santana, Valderi Reis Quietinho Leithardt

The paper proposes a unified, microservices-based platform that integrates endpoint telemetry and predictive NLP models to provide real-time, correlated alerting for security risks and hate speech.

View →