~ similar to 2606.27543· 14 results
Eight voice cloning models are benchmarked on five paralinguistic tasks, showing most preserve signal with modest degradation. Cloning English clinical speech into Japanese outperforms raw cross-lingu…
This paper proposes a graph framework using pretrained SSL embeddings for speech analysis in Amyotrophic Lateral Sclerosis (ALS) patients, achieving better results than validation baselines on the SAN…
This paper compares the performance of speech deepfake countermeasures using equal error rate (EER) and half total error rate (HTER) on different datasets. It also evaluates the effectiveness of popul…
Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close +1 more
The authors present ZONOS2 8B, a TTS model with improved naturalness, prosody, and voice cloning fidelity, achieved through scaling, data expansion, and simplification.
This paper evaluates the use of large audio language models for speaker verification systems against conventional pipelines and finds that task-specific adaptation improves performance.
David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer +10 more
The paper introduces the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech, speech-to-speech, speech understanding, and automatic speech recognition…
This paper proposes a novel pretraining method for Acoustic-to-Articulatory Inversion (AAI) using Phoneme Labels, Articulatory Feature Labels, and Critical-articulator Labels, improving performance an…
This paper conditions a diffusion-based speech enhancement model on wav2vec 2.0 features using Feature-wise Linear Modulation (FiLM), achieving competitive performance on VoiceBank-DEMAND and LibriMix…
Sicheng Yang, Shulan Ruan, Shiwei Wu, Yu Liu +3 more
PolySpeech-100 introduces a massive, multi-lingual benchmark covering 110 linguistic variants to rigorously test Speech-LLMs, demonstrating that open-source models struggle with low-resource languages…
A lightweight framework for automated pronunciation assessment using native speech resources and unsupervised or lightly calibrated methods.
The paper proposes methods for generating global prosodic embeddings using auto-encoder models of pitch and energy, demonstrating competitive or superior performance under challenging conditions.