20 results for “Zero-shot text-to-speech, Singlish, Fine-tuning, Accent similarity, Naturalness”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper fine-tunes two zero-shot text-to-speech models, Chatterbox and CosyVoice 3, on Singlish speakers from the IMDA National Speech Corpus to improve accent similarity and naturalness.
Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close +1 more
The authors present ZONOS2 8B, a TTS model with improved naturalness, prosody, and voice cloning fidelity, achieved through scaling, data expansion, and simplification.
Yujie Tu, Yifan Yang, Tianrui Wang, Yanqiao Zhu +32 more
The paper introduces GigaSpeechBench, a comprehensive multilingual and multidimensional ASR & AST benchmark with 680 hours of human-annotated speech, featuring 12 low-resource languages, 6 Chinese dia…
Sicheng Yang, Shulan Ruan, Shiwei Wu, Yu Liu +3 more
PolySpeech-100 introduces a massive, multi-lingual benchmark covering 110 linguistic variants to rigorously test Speech-LLMs, demonstrating that open-source models struggle with low-resource languages…
Sujith Pulikodan, Agneedh Basu, Pavan Kumar, Pranav Bhat +3 more
This paper investigates the effectiveness of incorporating synthetic speech data in Automatic Speech Recognition (ASR) Systems for three Indic languages by analyzing performance gains, script sources,…
SooHwan Eom, Hee Suk Yoon, Eunseop Yoon, Mark Hasegawa-Johnson +1 more
The paper proposes RTFree-F5, a method to make flow-matching TTS models like F5-TTS independent of reference transcripts, improving performance and naturalness for dysarthric speakers.
Haechan Kim, Seungjun Chung, Inkyu Park, Jihoo Lee +1 more
The paper introduces three new Korean speech benchmarks (KVoiceBench, KOpenAudioBench, and KMMAU) to evaluate SpeechLMs, demonstrating that English-centric evaluation fails to capture performance gaps…
The paper proposes Pitch-Accent-focused Speech Quality Assessment (PASQA) to explicitly target pitch-accent correctness in speech quality assessment, using a controlled Japanese accent-error dataset a…
This paper presents a method for building a compact Hindi text-to-speech model by pruning a large teacher model under a severe data budget, achieving state-of-the-art performance.
Shuai Wang, Zihan Qian, Ke Zhang, Jiangyu Han +8 more
This paper introduces the REAL-TSE Challenge, a satellite challenge on target speaker extraction from real conversational recordings, and describes its task definition, datasets, evaluation protocol,…
The paper introduces BEA-Dialogue+, an expanded 200-hour corpus for Hungarian conversational ASR, demonstrating that while larger data is challenging, specialized fine-tuning techniques significantly…
The paper describes a method to optimize the style vector for text-to-speech systems without a reference encoder, improving similarity and acceptance rate.
Sujith Pulikodan, Agneedh Basu, Saurabh Kumar, Pranav Bhat +4 more
The paper introduces a new inclusive, multimodal Hindi ASR benchmark with real-world recordings and diverse demographic groups, enabling more robust and realistic evaluation.
Yuxiang Zhao, Yichi Zhang, Yanjie An, Yanqiao Zhu +9 more
X-Translator is a low-cost modular system for real-time speech-to-speech translation, using streaming ASR, machine translation, and prompt-conditioned TTS, with session-level control.
Kaicheng Luo, Xuefei Gong, Yutao Sun, Jinling He +5 more
This paper introduces StellarTTS, a mobile-optimized non-autoregressive text-to-speech framework with sparse temporal embeddings and a semantic-aware codec, achieving lower latency and stronger robust…
Chatterbox-Flash introduces a prior-calibrated block diffusion model for zero-shot TTS that achieves high-fidelity, streaming synthesis with significantly lower computational overhead than existing me…
This paper introduces a pipeline to extract grammatical rules, example sentences, and lexicons from grammar books and generates synthetic parallel corpora for fine-tuning machine translation models on…