20 results for “vocal effort classification”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper introduces WavLM for vocal effort classification and improves performance through data augmentation and Gaussian-neighbor soft labels, achieving a new state-of-the-art on AVID.
Seolhee Lee, Minsu Kang, Yangsun Lee, Woosun Min +2 more
The paper introduces the Designed Vocalizations Dataset for AI-based voice conversion research on non-human vocalizations and effects, providing a standardized test set and benchmark results.
The paper introduces MMGenre, a benchmark for multi-genre singing voice synthesis diagnosis, revealing limited genre discrimination and proposing lightweight genre-specific continued training.
Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close +1 more
The authors present ZONOS2 8B, a TTS model with improved naturalness, prosody, and voice cloning fidelity, achieved through scaling, data expansion, and simplification.
The paper proposes a novel workflow to extract fine-grained regional accent features in Brazilian Portuguese using only acoustic labels and a phoneme-based forced aligner, showing that localized featu…
A Wave-U-Net model is trained to extract a fundamental waveform from input speech signals for accurate and robust instantaneous pitch estimation.
David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer +10 more
The paper introduces the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech, speech-to-speech, speech understanding, and automatic speech recognition…
The paper proposes Pitch-Accent-focused Speech Quality Assessment (PASQA) to explicitly target pitch-accent correctness in speech quality assessment, using a controlled Japanese accent-error dataset a…
This paper proposes a graph framework using pretrained SSL embeddings for speech analysis in Amyotrophic Lateral Sclerosis (ALS) patients, achieving better results than validation baselines on the SAN…
This paper introduces DG^VoiC, a voice clustering framework for customer verification and speaker linking in call-centre audio using sensitive information-aligned anonymisation, speech-focused preproc…
The study investigates the limitations of Phonetic Foundation Models (PFMs) for Speech-to-IPA transcription using Grapheme-to-Phoneme (G2P) labels and proposes a new approach based on continuous Artic…
This paper introduces CoughPhase-CLR, a self-supervised learning framework for cough representation learning using physiological phases, outperforming standard techniques on five downstream tasks.
The paper introduces Kiwano, an open-source toolkit for speaker verification research using PyTorch, offering standardized recipes, pretrained models, and integration of various architectures.
This paper introduces PolSeT, a Polish psychoacoustic and Music Information Retrieval dataset with 1901 descriptors and 18 instrument sound ratings.
A lightweight framework for automated pronunciation assessment using native speech resources and unsupervised or lightly calibrated methods.
This paper proposes a voice concept bottleneck framework for interpretable health assessment using an audio language model.
Yujie Tu, Yifan Yang, Tianrui Wang, Yanqiao Zhu +32 more
The paper introduces GigaSpeechBench, a comprehensive multilingual and multidimensional ASR & AST benchmark with 680 hours of human-annotated speech, featuring 12 low-resource languages, 6 Chinese dia…
Megan Wei, Deepali Aneja, Jiaqi Su, Yunyun Wang +2 more
This paper introduces MMEE, a multilingual and multi-emotion corpus for emphasis detection, and evaluates two state-of-the-art models under various settings.