~ similar to 2606.17662· 15 results
Sujith Pulikodan, Agneedh Basu, Saurabh Kumar, Pranav Bhat +4 more
The paper introduces a new inclusive, multimodal Hindi ASR benchmark with real-world recordings and diverse demographic groups, enabling more robust and realistic evaluation.
The paper introduces Script-Normalized WER (SN-WER), a novel evaluation metric that transliterates ASR transcripts into a canonical script to accurately measure speech recognition performance across d…
Eight voice cloning models are benchmarked on five paralinguistic tasks, showing most preserve signal with modest degradation. Cloning English clinical speech into Japanese outperforms raw cross-lingu…
This paper proposes a multimodal framework for jointly improving Automatic Speech Recognition (ASR) and Dialect Identification (DID) in Indian languages using a Bottleneck Encoder, RoBERTa encoder, ga…
Yujie Tu, Yifan Yang, Tianrui Wang, Yanqiao Zhu +32 more
The paper introduces GigaSpeechBench, a comprehensive multilingual and multidimensional ASR & AST benchmark with 680 hours of human-annotated speech, featuring 12 low-resource languages, 6 Chinese dia…
This paper introduces a pipeline to extract grammatical rules, example sentences, and lexicons from grammar books and generates synthetic parallel corpora for fine-tuning machine translation models on…
This paper explores the effect of conversational timing properties on automatic speech recognition (ASR) systems by controlling and optimizing pause and overlap timing distributions.
Ruchao Fan, Yiming Wang, Rui Zhao, Liliang Ren +9 more
This paper proposes Joint Speech-Text Interleaved Pretraining (JSTIP) for speech recognition, which constructs interleaved speech-text sequences and achieves consistent entity accuracy improvement.
This paper demonstrates that compact, domain-specialized Automatic Speech Recognition (ASR) models significantly outperform large, general-purpose foundation models for conversational speech across 19…
A lightweight framework for automated pronunciation assessment using native speech resources and unsupervised or lightly calibrated methods.
This paper presents a method for building a compact Hindi text-to-speech model by pruning a large teacher model under a severe data budget, achieving state-of-the-art performance.
Sicheng Yang, Shulan Ruan, Shiwei Wu, Yu Liu +3 more
PolySpeech-100 introduces a massive, multi-lingual benchmark covering 110 linguistic variants to rigorously test Speech-LLMs, demonstrating that open-source models struggle with low-resource languages…