20 results for “voice cloning”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Eight voice cloning models are benchmarked on five paralinguistic tasks, showing most preserve signal with modest degradation. Cloning English clinical speech into Japanese outperforms raw cross-lingu…
Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close +1 more
The authors present ZONOS2 8B, a TTS model with improved naturalness, prosody, and voice cloning fidelity, achieved through scaling, data expansion, and simplification.
Sujith Pulikodan, Agneedh Basu, Pavan Kumar, Pranav Bhat +3 more
This paper investigates the effectiveness of incorporating synthetic speech data in Automatic Speech Recognition (ASR) Systems for three Indic languages by analyzing performance gains, script sources,…
This paper evaluates the use of large audio language models for speaker verification systems against conventional pipelines and finds that task-specific adaptation improves performance.
The study tests the semantic independence of Resemble AI's deepfake audio detector, DETECT-3B-Omni, using 10,240 audio samples from various speakers and AI voice-cloning systems, and shows that the ac…
MindVoice is a neuro-to-speech framework that uses pretrained priors to disentangle and reconstruct intelligible speech from noisy, non-invasive neural signals, significantly outperforming existing me…
Seolhee Lee, Minsu Kang, Yangsun Lee, Woosun Min +2 more
The paper introduces the Designed Vocalizations Dataset for AI-based voice conversion research on non-human vocalizations and effects, providing a standardized test set and benchmark results.
The paper describes a method to optimize the style vector for text-to-speech systems without a reference encoder, improving similarity and acceptance rate.
SooHwan Eom, Hee Suk Yoon, Eunseop Yoon, Mark Hasegawa-Johnson +1 more
The paper proposes RTFree-F5, a method to make flow-matching TTS models like F5-TTS independent of reference transcripts, improving performance and naturalness for dysarthric speakers.
Yudong Li, Zihao Fang, Junwen Qiu, Ruihai Jing +3 more
This paper introduces Speaker Anonymization (SA) as a novel perturbation mechanism for zero-shot voice conversion, balancing timbre leakage and prosodic utility while enabling strictly causal, zero-lo…
Hanlei Zhang, Zhongming Ma, Mingyang Zhang, Tengfei Liu +2 more
The paper proposes TRIDENT, a framework to restore a source speaker's identity from converted audio using a three-pronged architecture.
Ye Lu, Yihan Yan, Zhaoyang Zhang, Zhitao Ou +3 more
This paper introduces Audio BERT (AuB) and SpInv, methods for recovering embeddings from speech tokens and performing speaker inversion attacks using only three seconds of frontend output.
This paper introduces SSTMark, a training-free speech watermarking framework that encodes watermark information into the semantic content of generated speech.
Yifan Liao, Yule Liu, Zhen Sun, Zongmin Zhang +4 more
The paper introduces MARS, a novel meta-adversarial framework that significantly improves black-box adversarial attacks against state-of-the-art Singing Voice Deepfake Detection (SVDD) systems by esca…
This paper introduces WanSong, a diffusion-based model for long-form, commercial-grade song generation that directly generates high-fidelity, multilingual songs up to 5 minutes and outputs dual stems,…
This paper fine-tunes two zero-shot text-to-speech models, Chatterbox and CosyVoice 3, on Singlish speakers from the IMDA National Speech Corpus to improve accent similarity and naturalness.
This paper introduces WavLM for vocal effort classification and improves performance through data augmentation and Gaussian-neighbor soft labels, achieving a new state-of-the-art on AVID.
Yuxiang Zhao, Yichi Zhang, Yanjie An, Yanqiao Zhu +9 more
X-Translator is a low-cost modular system for real-time speech-to-speech translation, using streaming ASR, machine translation, and prompt-conditioned TTS, with session-level control.
MelShield is a robust, in-generation audio watermarking framework that embeds identifiable signals into AI-generated speech in the Mel-spectrogram domain for reliable copyright protection and attribut…