~ similar to 2607.01747· 20 results
This paper explores the effect of different narrative strategies in data visualizations on eliciting prosocial feelings and behaviors during humanitarian crises.
The paper introduces ImputeViz, a visual analytics dashboard for diagnosing missing data, configuring imputation models, and evaluating results using various methods, including gKNN.
Abhraneel Sarma, Matthew Kay, Sheng Long, Michael Correll +1 more
This paper investigates the impact of performance-based financial incentives on task performance in visualization studies.
This paper introduces topological-geometrical metrics to estimate structural causal effects that are missed by traditional mean-based methods, proposing a new concept called topological ignorability.
This paper evaluates Centaur, a foundation model trained on psychological experiments, for program comprehension tasks and compares its performance to Llama 3.1.
Jiwon Kim, Maya Ajit, Sherry Gong, Soorya Ram Shimgekar +3 more
The paper introduces LLUMI, an open-source framework that improves LLM writing assistance for mental health support using community feedback, demonstrating comparable performance to proprietary models…
Jessica Woodhams, Amy Burrell, Wanyin Li, Fahim Ahmed +8 more
An industrial evaluation of an AI-enabled decision-support tool for crime linkage analysis was conducted, showing analysts selectively used AI predictions and valued model features.
The paper proposes a novel two-stage framework to differentially privatize tables of counts by focusing on preserving the accuracy of the underlying count distribution, introducing the specialized cyc…
The paper provides a formal statistical and conceptual framework for defining and measuring 'pairwise reference alignment,' which quantifies how well a model's scoring function agrees with a given ref…
Han Jiang, Sunbeom Kwon, Jinwen Luo, Ziang Xiao +1 more
This paper evaluates the reliability of using item response theory (IRT) models for AI benchmarking, comparing four estimation tools under various simulation conditions.
This study evaluated a personality-conditional cybersecurity training system, TailoredSec, finding that routing content based on a user's Five-Factor Model (FFM) trait significantly improved post-trai…
The paper introduces the Causal Sensitivity Score (CSS), an interventional metric that reveals that standard coverage-based evaluations fail to detect critical responsiveness deficits in clinical LLMs…
This study compares various authorship attribution methods on Japanese web reviews, finding that while BERT fine-tuning performs best, TF-IDF+LR offers superior stability and efficiency for large-scal…
The paper proposes MVRAF, a data-driven framework that quantifies vulnerability risk in large-scale cloud infrastructure by integrating multiple attack attributes and analyzing cumulative risk distrib…
Liuliu Chen, Elise R. Carrotte, Brian E. Chapman, Jo Robinson +1 more
The paper introduces FigSIM, the first fine-grained dataset for analyzing suicide memes, which is used to benchmark models across tasks like suicide severity and figurative language detection.
The paper introduces a 'replication-first' paradigm for LLM behavioral benchmarking, demonstrating that this rigorous approach uncovers significant, non-obvious performance drops between successive mo…