~ similar to 2608.13717· 17 results
Rongshen He, Xinyu Liang, Dekun Chen, Jiaqi Li +2 more
This paper introduces a training recipe for sentence-level and long-form streaming speech-to-speech translation using only 2k hours of paired cross-lingual data and auxiliary supervision.
Junlong Tong, Yao Zhang, Anhao Zhao, Yingqi Fan +2 more
ProactiveLLM introduces a novel framework that enables streaming LLMs to actively decide when to interact with incoming data by leveraging the model's internal states, significantly reducing latency w…
This paper proposes cumsum-composable phase transport, a streaming-native temporal layer for keyword spotting using unitary transport, prefix differences, and gated residual updates.
A lightweight framework for automated pronunciation assessment using native speech resources and unsupervised or lightly calibrated methods.
Sicheng Yang, Shulan Ruan, Shiwei Wu, Yu Liu +3 more
PolySpeech-100 introduces a massive, multi-lingual benchmark covering 110 linguistic variants to rigorously test Speech-LLMs, demonstrating that open-source models struggle with low-resource languages…
Chen Yang, Chufan Yu, Hanfu Chen, Jie Zhu +21 more
MOSS-Audio is a unified audio-language model designed for comprehensive understanding of speech, environmental sounds, and music, achieving strong performance across various audio-grounded tasks.
Ruchao Fan, Yiming Wang, Rui Zhao, Liliang Ren +9 more
This paper proposes Joint Speech-Text Interleaved Pretraining (JSTIP) for speech recognition, which constructs interleaved speech-text sequences and achieves consistent entity accuracy improvement.
Chatterbox-Flash introduces a prior-calibrated block diffusion model for zero-shot TTS that achieves high-fidelity, streaming synthesis with significantly lower computational overhead than existing me…
Cheng-Kang Chou, Ming-To Chuang, Ke-Han Lu, Chan-Jan Hsu +1 more
This paper investigates timestamp drift in modern autoregressive ASR systems and proposes REDDIT, a two-stage post-training framework to correct timestamps while avoiding forgetting.
Yuxiang Zhao, Yichi Zhang, Yanjie An, Yanqiao Zhu +9 more
X-Translator is a low-cost modular system for real-time speech-to-speech translation, using streaming ASR, machine translation, and prompt-conditioned TTS, with session-level control.
This paper uses continual learning with explicit disfluency tokens to improve Automatic Speech Recognition (ASR) systems on disfluent speech, addressing the information loss and hallucinations caused…
Youngwon Choi, Hyeonyu Kim, Taeyoun Kwon, Donghyuk Jung +1 more
CORTIS is a text-only adaptation framework that fine-tunes spoken language models for task-oriented voice agents using text-form task supervision.
Hanke Xie, Haopeng Lin, Jiale Qian, Dake Guo +12 more
This paper proposes SemBridge, a framework for continuous-latent autoregressive speech generation using discrete semantic tokens for supervision.