19 results for “speech-text interleaving”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper analyzes speech-text interleaved language models and reveals that they go through an implicit transcription phase in which spoken words become decodable as text in intermediate layers.
Ruchao Fan, Yiming Wang, Rui Zhao, Liliang Ren +9 more
This paper proposes Joint Speech-Text Interleaved Pretraining (JSTIP) for speech recognition, which constructs interleaved speech-text sequences and achieves consistent entity accuracy improvement.
Ke-Han Lu, Keqi Deng, Ruchao Fan, Rui Zhao +1 more
The paper proposes SpeechKV, a method to compress speech sequences inside large language models using a learned pooling, maintaining performance and delivering decoding speedup.
SooHwan Eom, Hee Suk Yoon, Eunseop Yoon, Mark Hasegawa-Johnson +1 more
The paper proposes RTFree-F5, a method to make flow-matching TTS models like F5-TTS independent of reference transcripts, improving performance and naturalness for dysarthric speakers.
Bohan Li, Shi Lian, Hankun Wang, Yiwei Guo +5 more
HoliTok introduces a novel continuous holistic tokenization model that provides a unified, high-fidelity latent representation for simultaneously supporting both speech generation and speech understan…
Sicheng Yang, Shulan Ruan, Shiwei Wu, Yu Liu +3 more
PolySpeech-100 introduces a massive, multi-lingual benchmark covering 110 linguistic variants to rigorously test Speech-LLMs, demonstrating that open-source models struggle with low-resource languages…
This paper presents a method for building a compact Hindi text-to-speech model by pruning a large teacher model under a severe data budget, achieving state-of-the-art performance.
This paper explores the effect of conversational timing properties on automatic speech recognition (ASR) systems by controlling and optimizing pause and overlap timing distributions.
The paper introduces BEA-Dialogue+, an expanded 200-hour corpus for Hungarian conversational ASR, demonstrating that while larger data is challenging, specialized fine-tuning techniques significantly…
A lightweight framework for automated pronunciation assessment using native speech resources and unsupervised or lightly calibrated methods.
Rongshen He, Xinyu Liang, Dekun Chen, Jiaqi Li +2 more
This paper introduces a training recipe for sentence-level and long-form streaming speech-to-speech translation using only 2k hours of paired cross-lingual data and auxiliary supervision.
Yujie Tu, Yifan Yang, Tianrui Wang, Yanqiao Zhu +32 more
The paper introduces GigaSpeechBench, a comprehensive multilingual and multidimensional ASR & AST benchmark with 680 hours of human-annotated speech, featuring 12 low-resource languages, 6 Chinese dia…
Jing Peng, Junhao Du, Chenghao Wang, Hanqi Li +20 more
The paper introduces SURE, a unified framework designed to standardize and improve the comparability and reproducibility of evaluations for advanced speech understanding models.
Murmur is an efficient inference system for long-form ASR that resolves the accuracy-latency trade-off by optimizing both inter-chunk processing and intra-chunk attention mechanisms.
Kaicheng Luo, Xuefei Gong, Yutao Sun, Jinling He +5 more
This paper introduces StellarTTS, a mobile-optimized non-autoregressive text-to-speech framework with sparse temporal embeddings and a semantic-aware codec, achieving lower latency and stronger robust…
Chen Yang, Chufan Yu, Hanfu Chen, Jie Zhu +21 more
MOSS-Audio is a unified audio-language model designed for comprehensive understanding of speech, environmental sounds, and music, achieving strong performance across various audio-grounded tasks.
Chatterbox-Flash introduces a prior-calibrated block diffusion model for zero-shot TTS that achieves high-fidelity, streaming synthesis with significantly lower computational overhead than existing me…