ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “data analysis”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.CLcs.AIEmpiricalRecentJul 7, 2026

Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities

So Hasegawa, Shailaja Keyur Sampat, Lei Liu, Wei-Peng Chen

The paper introduces DataGovBench, a benchmark for evaluating Large Language Models in real-world data analysis scenarios, revealing significant performance gaps with state-of-the-art models.

View →
cs.HCEmpiricalRecentJul 15, 2026

ZipLine: Visual Analysis of Multivariate Graphs with Predicate Logic

Sjoerd Vink, Suyang Li, Brian Montambault, Michael Behrisch +2 more

This paper introduces ZipLine, a system for integrative analysis of multivariate graphs through a unified predicate language and learning algorithm.

View →
cs.CRRecentMay 25, 2026

Efficient and Privacy-Preserving Distribution Statistics Analytics on Mobile Spatial Data

Xuhao Ren, Mingyang Zhao, Ruichen Zhang, Liehuang Zhu +1 more

The paper proposes eSpat-B and eSpat+ systems to enable efficient and privacy-preserving distribution statistics analysis on massive, dynamic mobile spatial data.

View →
cs.LGcs.AIcs.CLRecentMay 28, 2026

LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis

Kewei Xu, Xiaoben Lu, Shuofei Qiao, Zihan Ding +3 more

The paper introduces LongDS, a new benchmark for long-horizon, multi-turn data analysis, demonstrating that current AI agents struggle significantly with maintaining and updating complex analytical st…

View →
cs.LGcs.AIRecentMay 29, 2026

dashi: A Python library for Dataset Shift Characterization to Support Trustworthy AI Development and Deployment

David Fernández-Narro, Pablo Ferri, Ángel Sánchez-García, Juan M. García-Gómez +1 more

The paper introduces 'dashi,' an open-source Python library that provides comprehensive tools for characterizing dataset shifts (covariate, prior, concept) to ensure robust and trustworthy AI developm…

View →
cs.AIcs.DBRecentMay 27, 2026

A Query Engine for the Agents

Kenny Daniel

The paper introduces Hyperparam, a set of lightweight JavaScript libraries designed to enable direct, model-aware querying of unstructured data (like agent traces) within client-side AI applications.

View →
cs.CRcs.DBRecentApr 8, 2026

Interpreting the Error of Differentially Private Median Queries through Randomization Intervals

Thomas Humphries, Tim Li, Shufan Zhang, Karl Knopf +1 more

The paper introduces PostRI, a novel method that allows for computing a Randomization Interval (RI) for differentially private median queries after the median has already been estimated, significantly…

View →
cs.HCcs.AIcs.LGRecentMay 27, 2026

SmartIterator: Visual Analytics Workflows for Supervising Unsupervised Data Grouping

Gennady Andrienko, Natalia Andrienko

The paper introduces SmartIterator (SI), a visual analytics framework that systematically guides analysts through the complex process of evaluating and understanding how data groupings change across p…

View →
cs.HCcs.LGEmpiricalRecentJul 9, 2026

ImputeViz: A Visual Analytics Dashboard for Diagnosing Missing Data and Comparing Imputation Methods

Aitik Dandapat, Lalith Punepalle Raveendrareddy, Mithilesh Kumar Singh, Klaus Mueller

The paper introduces ImputeViz, a visual analytics dashboard for diagnosing missing data, configuring imputation models, and evaluating results using various methods, including gKNN.

View →
stat.MEmath.STstat.MLEmpiricalRecentJul 16, 2026

Post Hoc Inference for Component Attribution in Multivariate Change-Point Detection

Dhia-Elhaq Ouerfelli, Sylvain Arlot, Kevin Bleakley, Patrick Pamphile

This paper proposes statistical procedures to identify coordinates responsible for change-points in multivariate time series data.

View →
cs.CLSurveyRecentJul 3, 2026

Mental Health Disorder Detection Beyond Social Media: A Systematic Review of Available Datasets

Sadiya Sayara Chowdhury Puspo, Ana-Maria Bucur, Stevie Chancellor, Özlem Uzuner +1 more

This paper conducts a systematic review of non-social media, free-text datasets for mental health research, revealing their predominant focus on English and depression detection, and identifying key g…

View →
cs.SEcs.AIcs.LOEmpiricalRecentJul 4, 2026

Why3-py: A Tool for Formal Verification of Hypothesis Testing and Meta-Analysis in Python

Akira Tanaka, Yusuke Kawamoto

The paper proposes a formal verification framework for statistical programs in Python using Why3-py and extends StatWhy to verify meta-analysis methods.

View →
cs.CRcs.LGstat.CORecentMay 13, 2026

XAI and Statistical Analysis for Reliable Intrusion Detection in the UAVIDS-2025 Dataset: From Tree to Hybrid and Tabular DNN Ensembles

Iakovos-Christos Zarkadis, Christos Douligeris

This paper develops and analyzes various ensemble models, culminating in an XGBoost-based system, to reliably detect UAV intrusions using XAI and advanced statistical methods to pinpoint the root caus…

View →
cs.HCEmpiricalRecentJun 23, 2026

Optimizing Visual Analytics Workflows: From Theory to Practice

Philip Beaucamp, Alfie Abdul-Rahman, Rita Borgo, Wolfgang Jentner +4 more

This paper investigates ways to transform a theory-based methodology for optimizing visual analytics workflows from theory to practice using case studies.

View →
cs.AIcs.HCEmpiricalRecentJul 16, 2026

teLLMe Why (Ain't Nothing but a Jam): Exploratory Causal Analysis of Urban Driving Data

Qiwei Li, Jorge Ortiz

The paper presents teLLMe, a system for exploratory causal analysis of urban driving datasets using structured event tables, causal structure learning, and query-specific effect estimation.

View →
cs.CLcs.AIcs.CVRecentJun 4, 2026

Benchmarking Open-Source Layout Detection Models for Data Snapshot Extraction from Institutional Documents

AJ Carl P. Dy, Aivin V. Solatorio

This paper introduces a new benchmark dataset and evaluation framework for 'data snapshot extraction,' focusing on identifying and localizing semantically meaningful analytical artifacts within operat…

View →
cs.LGcs.AIcs.NERecentMay 28, 2026

Evolving Features vs Evolving Entire Trees with GP for Interpretable Survival Analysis

Thalea Schlender, Peter A. N. Bosman, Tanja Alderliesten

This paper proposes using genetic programming (GP) to jointly evolve both the feature sets and the structure of survival trees, resulting in highly interpretable and high-performing shallow models for…

View →
cs.DBcs.AIcs.CLEmpiricalRecentJun 26, 2026

Single and Multi Truth Data Fusion using Large Language Models

Hira Beril Kucuk, Norman W Paton, Jiaoyan Chen, Zhenyu Wu

This paper explores the use of Large Language Models (LLMs) in data fusion tasks for tabular data and shows their superiority over traditional methods.

View →
stat.MLcs.LGEmpiricalRecentJun 28, 2026

Gradient boosting with vector-valued leafs

David Cortes

This paper extends gradient boosting to functions of vector inputs using a simple algorithm with histogram-based decision trees.

View →