ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “vocal effort classification”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.SDcs.LGEmpiricalRecentJun 25, 2026

Advancing Speaker-Based Vocal Effort Classification with WavLM and Data Augmentation in Naturalistic Non-Calibrated Speech Recordings

Zahra Omidi, John H. L. Hansen

This paper introduces WavLM for vocal effort classification and improves performance through data augmentation and Gaussian-neighbor soft labels, achieving a new state-of-the-art on AVID.

View →
eess.AScs.SDDatasetRecentJul 23, 2026

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion

Seolhee Lee, Minsu Kang, Yangsun Lee, Woosun Min +2 more

The paper introduces the Designed Vocalizations Dataset for AI-based voice conversion research on non-human vocalizations and effects, providing a standardized test set and benchmark results.

View →
cs.SDEmpiricalRecentJul 8, 2026

MMGenre: Benchmarking Singing Voice Synthesis across Multiple Musical Genres

Wenhao Feng, Yuxun Tang, Jiatong Shi, Qin Jin

The paper introduces MMGenre, a benchmark for multi-genre singing voice synthesis diagnosis, revealing limited genre discrimination and proposing lightweight genre-specific continued training.

View →
cs.SDcs.AIEmpiricalRecentJun 23, 2026

ZONOS2 Technical Report

Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close +1 more

The authors present ZONOS2 8B, a TTS model with improved naturalness, prosody, and voice cloning fidelity, achieved through scaling, data expansion, and simplification.

View →
eess.AScs.CLRecentMay 28, 2026

Extracting accent features in spoken Brazilian Portuguese without sociolinguistic labels

Pedro H. L. Leite, Pedro Benevenuto Valadares, Luiz W. P. Biscainho

The paper proposes a novel workflow to extract fine-grained regional accent features in Brazilian Portuguese using only acoustic labels and a phoneme-based forced aligner, showing that localized featu…

View →
cs.SDEmpiricalRecentJun 12, 2026

Instantaneous Pitch Estimation via Wave-U-Net-Based Fundamental Waveform Enhancement

Junya Koguchi, Tomoki Koriyama

A Wave-U-Net model is trained to extract a fundamental waveform from input speech signals for accurate and robust instantaneous pitch estimation.

View →
cs.SDcs.AIEmpiricalRecentJul 16, 2026

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer +10 more

The paper introduces the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech, speech-to-speech, speech understanding, and automatic speech recognition…

View →
eess.AScs.CLcs.LGEmpiricalRecentJun 18, 2026

PASQA: Pitch-Accent-Focused Speech Quality Assessment Model Trained on Synthetic Speech with Accent Errors

Masaya Kawamura, Yuma Shirahata, Kentaro Mitsui, Reo Shimizu

The paper proposes Pitch-Accent-focused Speech Quality Assessment (PASQA) to explicitly target pitch-accent correctness in speech quality assessment, using a controlled Japanese accent-error dataset a…

View →
eess.AScs.SDEmpiricalRecentJul 28, 2026

Multi-Phonation Graph Learning with Self-Supervised Speech Embeddings for ALS Detection and Progression Prediction

Behrad TaghiBeyglou, Fatemeh Bagheri, Ervin Sejdic

This paper proposes a graph framework using pretrained SSL embeddings for speech analysis in Amyotrophic Lateral Sclerosis (ALS) patients, achieving better results than validation baselines on the SAN…

View →
cs.SDcs.AIcs.CLEmpiricalRecentJun 26, 2026

DG^VoiC: Speaker Clustering for Fraud Investigation under Real Call-Centre Conditions

Muhammad Shakeel Akram, Amal Htait, Abdul Hamid Sadka, Emma Meisingseth +1 more

This paper introduces DG^VoiC, a voice clustering framework for customer verification and speaker linking in call-centre audio using sensitive information-aligned anonymisation, speech-focused preproc…

View →
cs.SDeess.ASEmpiricalRecentJul 26, 2026

Improving Zero-Shot Phonetic Classification through Language-Agnostic Articulatory Features

Ryo Magoshi, Jaeyoung Lee, Shinsuke Sakai, Tatsuya Kawahara

The study investigates the limitations of Phonetic Foundation Models (PFMs) for Speech-to-IPA transcription using Grapheme-to-Phoneme (G2P) labels and proposes a new approach based on continuous Artic…

View →
cs.SDEmpiricalRecentJun 19, 2026

CoughPhase-CLR: Designing an acoustics-informed foundation model for coughing sound classification

Marius Moldovan, Anton Batliner, Thomas M. Berghaus, Björn W. Schuller +1 more

This paper introduces CoughPhase-CLR, a self-supervised learning framework for cough representation learning using physiological phases, outperforming standard techniques on five downstream tasks.

View →
cs.SDcs.LGEmpiricalRecentJun 21, 2026

Kiwano: A Cutting-Edge Open-Source Toolkit for Speaker Verification

Mickael Rouvier, Pierre Michel Bousquet

The paper introduces Kiwano, an open-source toolkit for speaker verification research using PyTorch, offering standardized recipes, pretrained models, and integration of various architectures.

View →
cs.SDeess.ASDatasetRecentJun 18, 2026

PolSeT: Polish Semantics of Timbre Dataset

Jan Jasiński

This paper introduces PolSeT, a Polish psychoacoustic and Music Information Retrieval dataset with 1901 descriptors and 18 instrument sound ratings.

View →
cs.CLcs.SDeess.ASEmpiricalRecentJun 18, 2026

Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal

Syeda Faiza Ahmed Sara, Shammur Absar Chowdhury

A lightweight framework for automated pronunciation assessment using native speech resources and unsupervised or lightly calibrated methods.

View →
eess.ASEmpiricalRecentJul 18, 2026

An Audio Language Model-Based Voice Concept Bottleneck Framework for Interpretable Health Assessment

Yu-Wen Chen, Julia Hirschberg

This paper proposes a voice concept bottleneck framework for interpretable health assessment using an audio language model.

View →
eess.ASEmpiricalRecentJun 27, 2026

GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

Yujie Tu, Yifan Yang, Tianrui Wang, Yanqiao Zhu +32 more

The paper introduces GigaSpeechBench, a comprehensive multilingual and multidimensional ASR & AST benchmark with 680 hours of human-annotated speech, featuring 12 low-resource languages, 6 Chinese dia…

View →
cs.CLcs.AIcs.LGEmpiricalRecentJun 26, 2026

Do Speech Emphasis Models Generalize across Languages and Emotions?

Megan Wei, Deepali Aneja, Jiaqi Su, Yunyun Wang +2 more

This paper introduces MMEE, a multilingual and multi-emotion corpus for emphasis detection, and evaluates two state-of-the-art models under various settings.

View →