20 results for “real-world data”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
The paper introduces DataGovBench, a benchmark for evaluating Large Language Models in real-world data analysis scenarios, revealing significant performance gaps with state-of-the-art models.
This paper introduces an open-source framework for evaluating the efficacy of AI agents powered by open-weight large language models on data preparation tasks in research using locally deployable mode…
This paper studies six metadata-generation settings for RDF datasets and evaluates their effectiveness and faithfulness in dataset search.
This study benchmarks four local LLMs for natural-language-to-SQL querying in biopharma manufacturing, finding that general-purpose code-tuned models like Llama 3.1 8B and Qwen 2.5 Coder 7B outperform…
Florian Schmid, Paul Primus, Alexander Fichtinger, Tara Jadidi +2 more
This paper introduces RealDESED, a new benchmark for domestic sound event detection with 5,710 recordings, precise annotations, and multi-annotator labeling.
The paper introduces Hyperparam, a set of lightweight JavaScript libraries designed to enable direct, model-aware querying of unstructured data (like agent traces) within client-side AI applications.
Zhensheng Wang, Xiaole Liu, Wenmian Yang, Kun Zhou +2 more
The paper introduces Open-Domain Tabular Question Answering for Future Data Forecasting and Reasoning, a new dataset and framework that enables LLMs to perform time-series forecasting and reasoning on…
The paper introduces Sophrosyne, a system that moderates LLM agent exploration in relational data systems, significantly reducing over-exploration and boosting SQL generation accuracy by guiding the a…
This study investigates how industrial practitioners perceive and manage data leakage in automotive perception systems, finding that leakage control is a socio-technical coordination problem requiring…
The paper introduces TaDaS, a framework that analyzes large-scale text archives to measure professional sentiment, finding that while AI discussion among economists is initially negative, the trend sh…
HuiMing Fan, Xiao Wang, Zheng Chu, Qianyu Wang +4 more
The paper argues that current search agents often verify existing knowledge rather than genuinely searching, and introduces LiveBrowseComp, a new benchmark to measure true evidence-driven discovery.
Kewei Xu, Xiaoben Lu, Shuofei Qiao, Zihan Ding +3 more
The paper introduces LongDS, a new benchmark for long-horizon, multi-turn data analysis, demonstrating that current AI agents struggle significantly with maintaining and updating complex analytical st…
Shuai Wang, Zihan Qian, Ke Zhang, Jiangyu Han +8 more
This paper introduces the REAL-TSE Challenge, a satellite challenge on target speaker extraction from real conversational recordings, and describes its task definition, datasets, evaluation protocol,…
This paper introduces a new benchmark dataset and evaluation framework for 'data snapshot extraction,' focusing on identifying and localizing semantically meaningful analytical artifacts within operat…
Gen Li, Jieyuan Lan, Pengcheng Xu, Zongyuan Wu +2 more
The paper presents GenWorld, an urban simulation infrastructure for LLM-agent studies, combining a synthetic city, agent-environment interface, and offline compilation of LLM-derived signals.
Yuchen He, Peizhi Ying, Liqi Cheng, Kuilin Peng +3 more
The paper builds a benchmark to evaluate the ability of multimodal large language models to extract accurate data tables from chart images, and proposes a human-centered approach to improve numerical…
The paper introduces RealVuln, a benchmark that demonstrates a clear three-tier performance hierarchy for security scanners on real-world code, with specialized tools significantly outperforming gener…
The paper proposes a decentralized, witnessing-zone architecture that enhances Proof-of-Location (PoL) to provide robust, auditable evidence of physical events, thereby improving sensor data trustwort…