ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “real-world data”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.CLcs.AIEmpiricalRecentJul 7, 2026

Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities

So Hasegawa, Shailaja Keyur Sampat, Lei Liu, Wei-Peng Chen

The paper introduces DataGovBench, a benchmark for evaluating Large Language Models in real-world data analysis scenarios, revealing significant performance gaps with state-of-the-art models.

View →
cs.AIcs.CLEmpiricalRecentJul 23, 2026

Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks

Mack Nixon, Liam Wright, Yevgeniya Kovalchuk, Alison Fang-Wei Wu +3 more

This paper introduces an open-source framework for evaluating the efficacy of AI agents powered by open-weight large language models on data preparation tasks in research using locally deployable mode…

View →
cs.IRcs.AIEmpiricalRecentJul 7, 2026

Faithful or Findable? Evaluating LLM-Generated Metadata for RDF Dataset Search

Riccardo Terrenzi, Serkan Ayvaz

This paper studies six metadata-generation settings for RDF datasets and evaluates their effectiveness and faithfulness in dataset search.

View →
cs.CLRecentMay 31, 2026

Benchmarking Local LLMs for Natural-Language-to-SQL Querying in Biopharmaceutical Manufacturing: An Empirical Benchmark on Consumer-Grade Hardware

Sagar Bhetwal, Rajan Bastakoti, Nirajan Acharya, Gaurav Kumar Gupta

This study benchmarks four local LLMs for natural-language-to-SQL querying in biopharma manufacturing, finding that general-purpose code-tuned models like Llama 3.1 8B and Qwen 2.5 Coder 7B outperform…

View →
eess.AScs.AIcs.SDDatasetRecentJul 18, 2026

RealDESED: A Real-World Domestic Sound Event Detection Benchmark

Florian Schmid, Paul Primus, Alexander Fichtinger, Tara Jadidi +2 more

This paper introduces RealDESED, a new benchmark for domestic sound event detection with 5,710 recordings, precise annotations, and multi-annotator labeling.

View →
cs.AIcs.DBRecentMay 27, 2026

A Query Engine for the Agents

Kenny Daniel

The paper introduces Hyperparam, a set of lightweight JavaScript libraries designed to enable direct, model-aware querying of unstructured data (like agent traces) within client-side AI applications.

View →
cs.IRcs.AIcs.CLRecentJun 1, 2026

ODTQA-FoRe: An Open-Domain Tabular Question Answering Dataset for Future Data Forecasting and Reasoning

Zhensheng Wang, Xiaole Liu, Wenmian Yang, Kun Zhou +2 more

The paper introduces Open-Domain Tabular Question Answering for Future Data Forecasting and Reasoning, a new dataset and framework that enables LLMs to perform time-series forecasting and reasoning on…

View →
cs.DBcs.AIRecentMay 29, 2026

Sophrosyne: Agentic Exploration of Relational Data Systems Needs Moderation

Madhav Jivrajani, Ramnatthan Alagappan, Aishwarya Ganesan

The paper introduces Sophrosyne, a system that moderates LLM agent exploration in relational data systems, significantly reducing over-exploration and boosting SQL generation accuracy by guiding the a…

View →
cs.CRcs.LGcs.SERecentApr 8, 2026

Data Leakage in Automotive Perception: Practitioners' Insights

Md Abu Ahammed Babu, Sushant Kumar Pandey, Darko Durisic, Andras Balint +1 more

This study investigates how industrial practitioners perceive and manage data leakage in automotive perception systems, finding that leakage control is a socio-technical coordination problem requiring…

View →
cs.CERecentJun 1, 2026

Are Economists Open to AI? Text as Data as Survey on Professional Sentiment and Academic Research Trends

Yi Wang, Lei Ge

The paper introduces TaDaS, a framework that analyzes large-scale text archives to measure professional sentiment, finding that while AI discussion among economists is initially negative, the trend sh…

View →
cs.AIRecentMay 27, 2026

LiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?

HuiMing Fan, Xiao Wang, Zheng Chu, Qianyu Wang +4 more

The paper argues that current search agents often verify existing knowledge rather than genuinely searching, and introduces LiveBrowseComp, a new benchmark to measure true evidence-driven discovery.

View →
cs.LGcs.AIcs.CLRecentMay 28, 2026

LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis

Kewei Xu, Xiaoben Lu, Shuofei Qiao, Zihan Ding +3 more

The paper introduces LongDS, a new benchmark for long-horizon, multi-turn data analysis, demonstrating that current AI agents struggle significantly with maintaining and updating complex analytical st…

View →
eess.AScs.SDEmpiricalRecentJul 16, 2026

SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

Shuai Wang, Zihan Qian, Ke Zhang, Jiangyu Han +8 more

This paper introduces the REAL-TSE Challenge, a satellite challenge on target speaker extraction from real conversational recordings, and describes its task definition, datasets, evaluation protocol,…

View →
cs.CLcs.AIcs.CVRecentJun 4, 2026

Benchmarking Open-Source Layout Detection Models for Data Snapshot Extraction from Institutional Documents

AJ Carl P. Dy, Aivin V. Solatorio

This paper introduces a new benchmark dataset and evaluation framework for 'data snapshot extraction,' focusing on identifying and localizing semantically meaningful analytical artifacts within operat…

View →
cs.MAEmpiricalRecentJun 26, 2026

GenWorld: Empirically Grounded Urban Simulation Infrastructure for Scalable LLM-Agent Studies

Gen Li, Jieyuan Lan, Pengcheng Xu, Zongyuan Wu +2 more

The paper presents GenWorld, an urban simulation infrastructure for LLM-agent studies, combining a synthetic city, agent-environment interface, and offline compilation of LLM-derived signals.

View →
cs.HCcs.AIEmpiricalRecentJun 29, 2026

Making Multimodal LLMs Reliable Chart Data Extractors: A Benchmark and Training Framework

Yuchen He, Peizhi Ying, Liqi Cheng, Kuilin Peng +3 more

The paper builds a benchmark to evaluate the ability of multimodal large language models to extract accurate data tables from chart images, and proposes a human-centered approach to improve numerical…

View →
cs.CRRecentApr 15, 2026

RealVuln: Benchmarking Rule-Based, General-Purpose LLM, and Security-Specialized Scanners on Real-World Code

John Pellew, Faizan Raza

The paper introduces RealVuln, a benchmark that demonstrates a clear three-tier performance hierarchy for security scanners on real-world code, with specialized tools significantly outperforming gener…

View →
cs.CRRecentMar 29, 2026

Decentralized Proof-of-Location for Content Provenance: Towards Capture-Time Authenticity

Eduardo Brito, Fernando Castillo, Amnir Hadachi, Ulrich Norbisrath +1 more

The paper proposes a decentralized, witnessing-zone architecture that enhances Proof-of-Location (PoL) to provide robust, auditable evidence of physical events, thereby improving sensor data trustwort…

View →