20 results for “data sources”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Haowen Wang, Yaxin Du, Jian Yang, Jiajun Wu +8 more
MIRA proposes a novel source-aware filtering framework that discovers and anchors evaluation rubrics during data selection, significantly improving code-oriented mid-training data quality while reduci…
This paper studies six metadata-generation settings for RDF datasets and evaluates their effectiveness and faithfulness in dataset search.
trasgoDP is an open-source Python framework for releasing tabular and location data under local differential privacy and geo-indistinguishability guarantees.
The study compares agentic data retrieval using unstructured web data versus structured, semantically-annotated datasets, concluding that semantic metadata remains essential for high-precision, reliab…
Jan Pennekamp, Johannes Lohmöller, David Schütte, Joscha Loos +1 more
This paper systematically analyzes 2.7 million arXiv submissions to demonstrate that nearly every preprint unintentionally discloses sensitive or unnecessary information through its source files, prop…
The paper introduces Hyperparam, a set of lightweight JavaScript libraries designed to enable direct, model-aware querying of unstructured data (like agent traces) within client-side AI applications.
The paper introduces DataGovBench, a benchmark for evaluating Large Language Models in real-world data analysis scenarios, revealing significant performance gaps with state-of-the-art models.
The paper introduces TaDaS, a framework that analyzes large-scale text archives to measure professional sentiment, finding that while AI discussion among economists is initially negative, the trend sh…
This paper introduces a new benchmark dataset and evaluation framework for 'data snapshot extraction,' focusing on identifying and localizing semantically meaningful analytical artifacts within operat…
The paper introduces a novel, scalable framework to monitor and classify dataset usage within research literature, addressing the current lack of infrastructure for tracking data citations.
Jailbreak bypasses database engines by reading storage files directly and materializing data as in-memory columnar buffers using LLM-assisted code synthesis.
The paper introduces 'dashi,' an open-source Python library that provides comprehensive tools for characterizing dataset shifts (covariate, prior, concept) to ensure robust and trustworthy AI developm…
The paper introduces FOSSIL, a new multilingual dataset and specialized workflow designed to significantly improve the extraction of citations embedded within complex footnotes common in law and human…
Jinheon Baek, Soyeong Jeong, Sangwoo Park, Woongyeong Yeo +4 more
OmniRetrieval introduces a unified framework that handles natural language queries across diverse, heterogeneous knowledge sources (text, relational, graphs) by dispatching source-native queries witho…
This paper introduces a framework to audit source-dependence in multi-source RAG systems, demonstrating that disagreement across institutional sources is a common and critical failure mode that curren…
The paper introduces Sophrosyne, a system that moderates LLM agent exploration in relational data systems, significantly reducing over-exploration and boosting SQL generation accuracy by guiding the a…
This paper introduces an open-source framework for evaluating the efficacy of AI agents powered by open-weight large language models on data preparation tasks in research using locally deployable mode…
Thomas Humphries, Tim Li, Shufan Zhang, Karl Knopf +1 more
The paper introduces PostRI, a novel method that allows for computing a Randomization Interval (RI) for differentially private median queries after the median has already been estimated, significantly…