ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Dataset”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.LGcs.AIRecentMay 28, 2026

idSCD: Identifying Training Datasets through Semantic Correlation Descriptors

Andrada Gobeaja, Ionut Hodoroaga, Elena Burceanu, Marius Leordeanu

The paper introduces a novel semantic fingerprinting approach using Semantic Correlation Descriptors (SCDs) to identify which specific datasets were used to train a model, demonstrating superior perfo…

View →
cs.SDeess.ASDatasetRecentJun 18, 2026

PolSeT: Polish Semantics of Timbre Dataset

Jan Jasiński

This paper introduces PolSeT, a Polish psychoacoustic and Music Information Retrieval dataset with 1901 descriptors and 18 instrument sound ratings.

View →
cs.CLRecentMay 28, 2026

AI for Monitoring and Classifying Data Used in Research Literature

Rafael Macalaba, Aivin V. Solatorio

The paper introduces a novel, scalable framework to monitor and classify dataset usage within research literature, addressing the current lack of infrastructure for tracking data citations.

View →
cs.CLcs.AIRecentMay 27, 2026

IPO-Mine: A Toolkit and Dataset for Section-Structured Analysis of Long, Multimodal IPO Documents

Michael Galarnyk, Siddharth Lohani, Vidhyakshaya Kannan, Sagnik Nandi +7 more

The paper introduces IPO-Mine, a comprehensive toolkit and large-scale dataset designed to enable standardized, multimodal analysis of extremely long and structurally complex Initial Public Offering (…

View →
cs.CLcs.AIcs.CVRecentJun 4, 2026

Benchmarking Open-Source Layout Detection Models for Data Snapshot Extraction from Institutional Documents

AJ Carl P. Dy, Aivin V. Solatorio

This paper introduces a new benchmark dataset and evaluation framework for 'data snapshot extraction,' focusing on identifying and localizing semantically meaningful analytical artifacts within operat…

View →
eess.AScs.SDDatasetRecentJul 23, 2026

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion

Seolhee Lee, Minsu Kang, Yangsun Lee, Woosun Min +2 more

The paper introduces the Designed Vocalizations Dataset for AI-based voice conversion research on non-human vocalizations and effects, providing a standardized test set and benchmark results.

View →
cs.CRRecentMay 27, 2026

Cybersecurity AI (CAI) Dataset

Víctor Mayoral-Vilches

The paper introduces the CAI Dataset, a massive, multi-terabyte corpus of real-world, hands-on cybersecurity LLM trajectories, designed to address the performance bottleneck caused by expert operator…

View →
cs.LGcs.AIRecentMay 29, 2026

dashi: A Python library for Dataset Shift Characterization to Support Trustworthy AI Development and Deployment

David Fernández-Narro, Pablo Ferri, Ángel Sánchez-García, Juan M. García-Gómez +1 more

The paper introduces 'dashi,' an open-source Python library that provides comprehensive tools for characterizing dataset shifts (covariate, prior, concept) to ensure robust and trustworthy AI developm…

View →
cs.CLSurveyRecentJul 3, 2026

Mental Health Disorder Detection Beyond Social Media: A Systematic Review of Available Datasets

Sadiya Sayara Chowdhury Puspo, Ana-Maria Bucur, Stevie Chancellor, Özlem Uzuner +1 more

This paper conducts a systematic review of non-social media, free-text datasets for mental health research, revealing their predominant focus on English and depression detection, and identifying key g…

View →
cs.PLEmpiricalRecentJul 21, 2026

VirtualSet: Typed Ontology Worlds as an LLM Generation Target for Grounded Queries and Guarded Decisions

Qunhui Zhang

This paper introduces VirtualSet, a live, receiver-typed ontology-world interface for large language models to interact with enterprise data, ensuring pre-execution semantics for guarded decisions.

View →
cs.LGcs.AIRecentMay 29, 2026

From Rashomon Theory to PRAXIS: Efficient Decision Tree Rashomon Sets

Zakk Heile, Hayden McTavish, Varun Babbar, Margo Seltzer +1 more

The paper introduces PRAXIS, a novel algorithm that efficiently approximates the computation of 'Rashomon sets' for decision trees, significantly reducing memory and runtime complexity.

View →
cs.LGcs.AIcs.CRRecentMay 19, 2026

LLM Benchmark Datasets Should Be Contamination-Resistant

Ali Al-Lawati, Jason Lucas, Dongwon Lee, Suhang Wang

The paper argues that current LLM benchmark datasets are often contaminated by being included in pretraining data, and proposes that future benchmarks must be contamination-resistant and support infer…

View →
cs.SDeess.ASeess.SPEmpiricalRecentJun 27, 2026

Underwater Source Detection and Classification for Signal-based Surveillance: Audio Dataset Curation and Cross-Domain Evaluation

Quoc Thinh Vo, David K. Han

This paper introduces a curated underwater audio dataset and proposes a margin-enhanced loss with feature alignment for underwater acoustic classification.

View →
cs.CVRecentJun 1, 2026

Places in the Wild: A Large, High-Resolution RAW Photograph Dataset for Ecologically Valid Vision Research

Michelle R. Greene

Places in the Wild introduces a massive, high-resolution RAW photograph dataset of 67,574 images captured in situ across 810 locations, providing unprecedented detail for ecologically valid vision res…

View →
cs.CLcs.AIEmpiricalRecentJul 7, 2026

Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities

So Hasegawa, Shailaja Keyur Sampat, Lei Liu, Wei-Peng Chen

The paper introduces DataGovBench, a benchmark for evaluating Large Language Models in real-world data analysis scenarios, revealing significant performance gaps with state-of-the-art models.

View →
cs.CLcs.AIEmpiricalRecentJul 27, 2026

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data

Zhen Huang, Yikun Wang, Shijie Xia, Pengfei Liu

The paper introduces DataOrchestra, a framework for example-specific processing of pretraining data for Large Language Models, achieving stable gains and reducing compute.

View →
cs.CLcs.AIRecentMay 28, 2026

MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settings

Valentina Bui Muti, Eugénie Dulout, Ziquan Fu

The paper introduces MedCase-Structured, a synthetic, FHIR-formatted dataset designed to benchmark diagnostic reasoning in realistic EHR settings, showing that LLMs perform worse on structured data th…

View →
cs.IRcs.AIEmpiricalRecentJul 7, 2026

Faithful or Findable? Evaluating LLM-Generated Metadata for RDF Dataset Search

Riccardo Terrenzi, Serkan Ayvaz

This paper studies six metadata-generation settings for RDF datasets and evaluates their effectiveness and faithfulness in dataset search.

View →