ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

~ similar to 2607.01747· 20 results

cs.HCEmpiricalRecentJul 3, 2026

Evaluating Affective Objectives: Statistical Numbing in Data Visualization

Elsie Lee-Robbins, Eytan Adar

This paper explores the effect of different narrative strategies in data visualizations on eliciting prosocial feelings and behaviors during humanitarian crises.

View →
cs.HCcs.LGEmpiricalRecentJul 9, 2026

ImputeViz: A Visual Analytics Dashboard for Diagnosing Missing Data and Comparing Imputation Methods

Aitik Dandapat, Lalith Punepalle Raveendrareddy, Mithilesh Kumar Singh, Klaus Mueller

The paper introduces ImputeViz, a visual analytics dashboard for diagnosing missing data, configuring imputation models, and evaluating results using various methods, including gKNN.

View →
cs.HCEmpiricalRecentJul 8, 2026

Should We Dangle a Carrot? The Effect of Performance-based Incentives in Visualization Experiments

Abhraneel Sarma, Matthew Kay, Sheng Long, Michael Correll +1 more

This paper investigates the impact of performance-based financial incentives on task performance in visualization studies.

View →
stat.MEcs.AIRecentMay 31, 2026

Topological Ignorability for Structural Causal Effects Beyond Means

Usef Faghihi

This paper introduces topological-geometrical metrics to estimate structural causal effects that are missed by traditional mean-based methods, proposing a new concept called topological ignorability.

View →
cs.SEEmpiricalRecentJul 13, 2026

Predicting Program Comprehension with Foundation Models of Human Cognition

Yannick Lehmen, Marvin Wyrich, Anna-Maria Maurer, Norman Peitek +1 more

This paper evaluates Centaur, a foundation model trained on psychological experiments, for program comprehension tasks and compares its performance to Llama 3.1.

View →
cs.HCcs.AIcs.CLRecentMay 28, 2026

LLUMI: Improving LLM Writing Assistance for Mental Health Support with Online Community Feedback

Jiwon Kim, Maya Ajit, Sherry Gong, Soorya Ram Shimgekar +3 more

The paper introduces LLUMI, an open-source framework that improves LLM writing assistance for mental health support using community feedback, demonstrating comparable performance to proprietary models…

View →
cs.HCcs.SEEmpiricalRecentJul 9, 2026

How Analysts Use AI in High-Stakes Crime Linkage: An Industrial Study

Jessica Woodhams, Amy Burrell, Wanyin Li, Fahim Ahmed +8 more

An industrial evaluation of an AI-enabled decision-support tool for crime linkage analysis was conducted, showing analysts selectively used AI predictions and valued model features.

View →
cs.CRRecentApr 1, 2026

Preserving Target Distributions With Differentially Private Count Mechanisms

Nitin Kohli, Paul Laskowski

The paper proposes a novel two-stage framework to differentially privatize tables of counts by focusing on preserving the accuracy of the underlying count distribution, introducing the specialized cyc…

View →
cs.CLcs.LGRecentMay 29, 2026

Pairwise Reference Alignment as a Model-Level Ordinal Observable

Mujing Li

The paper provides a formal statistical and conceptual framework for defining and measuring 'pairwise reference alignment,' which quantifies how well a model's scoring function agrees with a given ref…

View →
cs.AIEmpiricalRecentJul 16, 2026

Can We Trust Item Response Theory for AI Evaluation?

Han Jiang, Sunbeom Kwon, Jinwen Luo, Ziang Xiao +1 more

This paper evaluates the reliability of using item response theory (IRT) models for AI benchmarking, comparing four estimation tools under various simulation conditions.

View →
cs.CRcs.HCRecentMay 23, 2026

Routing Cybersecurity Awareness Training by FFM Personality Trait: A Quasi-Experimental Evaluation

Glory Okwata, Mohammad A. Razzaque

This study evaluated a personality-conditional cybersecurity training system, TailoredSec, finding that routing content based on a user's Five-Factor Model (FFM) trait significantly improved post-trai…

View →
cs.LGcs.AIcs.CLRecentMay 28, 2026

Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents

Matt Turk

The paper introduces the Causal Sensitivity Score (CSS), an interventional metric that reveals that standard coverage-based evaluations fail to detect critical responsiveness deficits in clinical LLMs…

View →
cs.CLcs.CRRecentMar 24, 2026

Foundational Study on Authorship Attribution of Japanese Web Reviews for Actor Analysis

Hiroshi Matsubara, Shingo Matsugaya, Taichi Aoki, Masaki Hashimoto

This study compares various authorship attribution methods on Japanese web reviews, finding that while BERT fine-tuning performs best, TF-IDF+LR offers superior stability and efficiency for large-scal…

View →
cs.CRRecentMar 30, 2026

Policy-Driven Vulnerability Risk Quantification framework for Large-Scale Cloud Infrastructure Data Security

Wanru Shao

The paper proposes MVRAF, a data-driven framework that quantifies vulnerability risk in large-scale cloud infrastructure by integrating multiple attack attributes and analyzing cumulative risk distrib…

View →
cs.CLcs.CVcs.CYRecentJun 1, 2026

FigSIM: A Dataset for Fine-grained Suicide Severity and Figurative Language in Suicide Memes

Liuliu Chen, Elise R. Carrotte, Brian E. Chapman, Jo Robinson +1 more

The paper introduces FigSIM, the first fine-grained dataset for analyzing suicide memes, which is used to benchmark models across tasks like suicide severity and figurative language detection.

View →
cs.CLcs.AIRecentMay 27, 2026

Let the Results Speak: A Replication-First Paradigm for LLM Behavioral Benchmarking

Yuming, Huang, Yao Liu, Lei Wang +1 more

The paper introduces a 'replication-first' paradigm for LLM behavioral benchmarking, demonstrating that this rigorous approach uncovers significant, non-obvious performance drops between successive mo…

View →