20 results for “datasets”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper conducts a systematic review of non-social media, free-text datasets for mental health research, revealing their predominant focus on English and depression detection, and identifying key g…
The paper argues that current LLM benchmark datasets are often contaminated by being included in pretraining data, and proposes that future benchmarks must be contamination-resistant and support infer…
This paper introduces PolSeT, a Polish psychoacoustic and Music Information Retrieval dataset with 1901 descriptors and 18 instrument sound ratings.
The paper introduces a novel semantic fingerprinting approach using Semantic Correlation Descriptors (SCDs) to identify which specific datasets were used to train a model, demonstrating superior perfo…
The paper introduces 'dashi,' an open-source Python library that provides comprehensive tools for characterizing dataset shifts (covariate, prior, concept) to ensure robust and trustworthy AI developm…
The paper introduces a novel, scalable framework to monitor and classify dataset usage within research literature, addressing the current lack of infrastructure for tracking data citations.
This paper introduces a new benchmark dataset and evaluation framework for 'data snapshot extraction,' focusing on identifying and localizing semantically meaningful analytical artifacts within operat…
The paper introduces IPO-Mine, a comprehensive toolkit and large-scale dataset designed to enable standardized, multimodal analysis of extremely long and structurally complex Initial Public Offering (…
Places in the Wild introduces a massive, high-resolution RAW photograph dataset of 67,574 images captured in situ across 810 locations, providing unprecedented detail for ecologically valid vision res…
The paper introduces DataGovBench, a benchmark for evaluating Large Language Models in real-world data analysis scenarios, revealing significant performance gaps with state-of-the-art models.
This paper studies six metadata-generation settings for RDF datasets and evaluates their effectiveness and faithfulness in dataset search.
This paper introduces VirtualSet, a live, receiver-typed ontology-world interface for large language models to interact with enterprise data, ensuring pre-execution semantics for guarded decisions.
Seolhee Lee, Minsu Kang, Yangsun Lee, Woosun Min +2 more
The paper introduces the Designed Vocalizations Dataset for AI-based voice conversion research on non-human vocalizations and effects, providing a standardized test set and benchmark results.
Zakk Heile, Hayden McTavish, Varun Babbar, Margo Seltzer +1 more
The paper introduces PRAXIS, a novel algorithm that efficiently approximates the computation of 'Rashomon sets' for decision trees, significantly reducing memory and runtime complexity.
The paper introduces the CAI Dataset, a massive, multi-terabyte corpus of real-world, hands-on cybersecurity LLM trajectories, designed to address the performance bottleneck caused by expert operator…
The paper introduces MedCase-Structured, a synthetic, FHIR-formatted dataset designed to benchmark diagnostic reasoning in realistic EHR settings, showing that LLMs perform worse on structured data th…