~ similar to 2606.30312· 20 results
Jiwon Kim, Maya Ajit, Sherry Gong, Soorya Ram Shimgekar +3 more
The paper introduces LLUMI, an open-source framework that improves LLM writing assistance for mental health support using community feedback, demonstrating comparable performance to proprietary models…
The paper introduces LinguIUTics, a system that significantly improves the classification of rare psychological defense mechanisms in conversational text by fine-tuning Qwen3-8B using specialized imba…
RealityTest introduces a large-scale, multimodal, and multilingual benchmark using real-world human data to test how AI systems disclose their identity, finding that context and phrasing are more crit…
This paper systematically measured web tracking across 20 popular AI chatbots, finding that a majority share both conversational content and user identity information with third parties.
This paper evaluates multiple LLMs (DeepSeek-R1, OpenBioLLM-Llama3, Qwen 3.5) for generating privacy-safe, high-quality synthetic mental health reports, demonstrating their effectiveness in expanding…
João Matos, Olivia Buege, Donny Cheung, Gary S. Collins +7 more
This paper analyzes 2,053 real patient-chatbot conversations and develops a patient simulator to evaluate the performance of LLMs in symptom assessment. Communication style was found to significantly…
This paper conducts a systematic review of non-social media, free-text datasets for mental health research, revealing their predominant focus on English and depression detection, and identifying key g…
The study tests the semantic independence of Resemble AI's deepfake audio detector, DETECT-3B-Omni, using 10,240 audio samples from various speakers and AI voice-cloning systems, and shows that the ac…
This paper proposes a multimodal cognitive impairment detection framework using large language models that integrates speech audio and transcripts, achieving a high classification accuracy.
Shuai Xiao, Su Liu, Weikai Zhou, Jialun Wu +3 more
Persona prompting does not universally improve LLM performance; instead, it systematically trades increased expertise depth for reduced clarity, making multi-metric evaluation essential.
The paper introduces WebPII, a novel, large-scale synthetic benchmark for detecting personally identifiable information (PII) in web screenshots, and demonstrates a model (WebRedact) that significantl…
Shashie Dilhara Batan Arachchige, Hassan Jameel Asghar, Benjamin Zi Hao Zhao, Dinusha Vatsalan +1 more
The paper proposes a character-level differential privacy mechanism to sanitize sensitive user prompts for LLMs, achieving high privacy for PII while maintaining utility for non-sensitive context.
The paper proposes a novel, locally deployable agentic workflow using large language models (LLMs) to accurately and privately detect various types of personally identifiable information (PII) within…
This paper investigates how the final prompt in conversational AI-search evaluations differs from the conversation history, using two corpora of commercial and PRISM conversations.
Michele Miranda, Xinlan Yan, Nishant Mishra, Rachel Murphy +3 more
This paper conducts the first comparative study of Differential Privacy (DP), Named Entity Recognition (NER), and Large Language Models (LLMs) for de-identifying Dutch clinical notes, finding that com…
This paper demonstrates that patient-facing RAG chatbots frequently expose sensitive system configurations, knowledge base details, and conversation history through client-server communication, posing…
The paper benchmarks local, offline LLMs for confidential translation workflows, demonstrating that while they are viable for privacy-sensitive use, they generally lag behind top commercial NMT system…