20 results for “data analysis”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
The paper introduces DataGovBench, a benchmark for evaluating Large Language Models in real-world data analysis scenarios, revealing significant performance gaps with state-of-the-art models.
Sjoerd Vink, Suyang Li, Brian Montambault, Michael Behrisch +2 more
This paper introduces ZipLine, a system for integrative analysis of multivariate graphs through a unified predicate language and learning algorithm.
Xuhao Ren, Mingyang Zhao, Ruichen Zhang, Liehuang Zhu +1 more
The paper proposes eSpat-B and eSpat+ systems to enable efficient and privacy-preserving distribution statistics analysis on massive, dynamic mobile spatial data.
Kewei Xu, Xiaoben Lu, Shuofei Qiao, Zihan Ding +3 more
The paper introduces LongDS, a new benchmark for long-horizon, multi-turn data analysis, demonstrating that current AI agents struggle significantly with maintaining and updating complex analytical st…
The paper introduces 'dashi,' an open-source Python library that provides comprehensive tools for characterizing dataset shifts (covariate, prior, concept) to ensure robust and trustworthy AI developm…
The paper introduces Hyperparam, a set of lightweight JavaScript libraries designed to enable direct, model-aware querying of unstructured data (like agent traces) within client-side AI applications.
Thomas Humphries, Tim Li, Shufan Zhang, Karl Knopf +1 more
The paper introduces PostRI, a novel method that allows for computing a Randomization Interval (RI) for differentially private median queries after the median has already been estimated, significantly…
The paper introduces SmartIterator (SI), a visual analytics framework that systematically guides analysts through the complex process of evaluating and understanding how data groupings change across p…
The paper introduces ImputeViz, a visual analytics dashboard for diagnosing missing data, configuring imputation models, and evaluating results using various methods, including gKNN.
This paper proposes statistical procedures to identify coordinates responsible for change-points in multivariate time series data.
This paper conducts a systematic review of non-social media, free-text datasets for mental health research, revealing their predominant focus on English and depression detection, and identifying key g…
The paper proposes a formal verification framework for statistical programs in Python using Why3-py and extends StatWhy to verify meta-analysis methods.
This paper develops and analyzes various ensemble models, culminating in an XGBoost-based system, to reliably detect UAV intrusions using XAI and advanced statistical methods to pinpoint the root caus…
This paper investigates ways to transform a theory-based methodology for optimizing visual analytics workflows from theory to practice using case studies.
The paper presents teLLMe, a system for exploratory causal analysis of urban driving datasets using structured event tables, causal structure learning, and query-specific effect estimation.
This paper introduces a new benchmark dataset and evaluation framework for 'data snapshot extraction,' focusing on identifying and localizing semantically meaningful analytical artifacts within operat…
This paper proposes using genetic programming (GP) to jointly evolve both the feature sets and the structure of survival trees, resulting in highly interpretable and high-performing shallow models for…
This paper explores the use of Large Language Models (LLMs) in data fusion tasks for tabular data and shows their superiority over traditional methods.
This paper extends gradient boosting to functions of vector inputs using a simple algorithm with histogram-based decision trees.