20 results for “data augmentation”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper proposes an uncertainty-guided synthetic context augmentation strategy for semantic segmentation models to improve performance on complex datasets with data sparsity and rare or visually di…
This paper explores the impact of poisoning attacks on data augmentation techniques, specifically on 3D point cloud datasets using GANs, and demonstrates that poisoning can evade augmentation and prop…
Xiaojing Chen, Jingqi Cheng, Xu Zhao, Wan Jiang +1 more
The paper introduces Score-Guided Classification (SGC), a novel framework that uses an unsupervised anomaly score as a 'Pathological Prior' to guide EEG-based depression detection, overcoming the limi…
Eight voice cloning models are benchmarked on five paralinguistic tasks, showing most preserve signal with modest degradation. Cloning English clinical speech into Japanese outperforms raw cross-lingu…
Ni Li, Nuohao Liu, Ryan Jacobs, Ajay Annamareddy +4 more
The paper proposes using a mask-conditioned latent diffusion model to generate synthetic, labeled TEM images for data augmentation, achieving small but measurable performance improvements in defect de…
This paper evaluates multiple LLMs (DeepSeek-R1, OpenBioLLM-Llama3, Qwen 3.5) for generating privacy-safe, high-quality synthetic mental health reports, demonstrating their effectiveness in expanding…
This paper explores the use of Large Language Models (LLMs) in data fusion tasks for tabular data and shows their superiority over traditional methods.
This paper introduces AutoBackSwap, a method to reduce reliance of classifiers on spurious backgrounds in image classification tasks using a secondary network and infilling.
Sujith Pulikodan, Agneedh Basu, Pavan Kumar, Pranav Bhat +3 more
This paper investigates the effectiveness of incorporating synthetic speech data in Automatic Speech Recognition (ASR) Systems for three Indic languages by analyzing performance gains, script sources,…
Weak self-training on synthetic data can amplify a language model's existing capabilities, but this effect is strictly dependent on the compatibility between the source and student models, not on the…
The paper argues that the standard FID metric is unreliable because its performance depends significantly on the geometric structure and density of the reference dataset, not just the sample quality.
This paper introduces a new benchmark dataset and evaluation framework for 'data snapshot extraction,' focusing on identifying and localizing semantically meaningful analytical artifacts within operat…
The paper demonstrates that using synthetic hand images containing accessories, generated via inpainting, significantly improves the robustness of hand detectors for safety-critical applications by cl…
The paper demonstrates that off-the-shelf image diffusion models, like Stable Diffusion, can be repurposed to generate synthetic structured data, posing a threat of ground truth drift in closed eviden…
Introduces Expanding Generative Flows (EFlows) and Expanding Flow Maps (EFMs) for generating outputs of varying sizes in continuous and discrete state spaces.
The paper proposes Attack-AAIRS, a novel framework that uses GAN-generated synthetic adversarial samples to enhance the robustness of skeleton-based person identification models against unseen attacks…
TabChange proposes a novel framework to generate natural and minimally altered counterfactual instances in tabular data by precisely controlling attribute modifications based on their relationship str…