ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

~ similar to 2608.13717· 17 results

cs.SDEmpiricalRecentJul 22, 2026

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision

Rongshen He, Xinyu Liang, Dekun Chen, Jiaqi Li +2 more

This paper introduces a training recipe for sentence-level and long-form streaming speech-to-speech translation using only 2k hours of paired cross-lingual data and auxiliary supervision.

View →
cs.CLRecentMay 30, 2026

ProactiveLLM: Learning Active Interaction for Streaming Large Language Models

Junlong Tong, Yao Zhang, Anhao Zhao, Yingqi Fan +2 more

ProactiveLLM introduces a novel framework that enables streaming LLMs to actively decide when to interact with incoming data by leveraging the model's internal states, significantly reducing latency w…

View →
cs.SDcs.LGEmpiricalRecentJul 22, 2026

Cumsum-Composable Phase Transport for Low-Cost Streaming Keyword Spotting

Mahesh Godavarti

This paper proposes cumsum-composable phase transport, a streaming-native temporal layer for keyword spotting using unitary transport, prefix differences, and gated residual updates.

View →
cs.CLcs.SDeess.ASEmpiricalRecentJun 18, 2026

Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal

Syeda Faiza Ahmed Sara, Shammur Absar Chowdhury

A lightweight framework for automated pronunciation assessment using native speech resources and unsupervised or lightly calibrated methods.

View →
cs.CLcs.AIeess.ASRecentMay 31, 2026

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects

Sicheng Yang, Shulan Ruan, Shiwei Wu, Yu Liu +3 more

PolySpeech-100 introduces a massive, multi-lingual benchmark covering 110 linguistic variants to rigorously test Speech-LLMs, demonstrating that open-source models struggle with low-resource languages…

View →
cs.SDcs.AIRecentJun 1, 2026

MOSS-Audio Technical Report

Chen Yang, Chufan Yu, Hanfu Chen, Jie Zhu +21 more

MOSS-Audio is a unified audio-language model designed for comprehensive understanding of speech, environmental sounds, and music, achieving strong performance across various audio-grounded tasks.

View →
cs.CLeess.ASEmpiricalRecentJul 2, 2026

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving

Ruchao Fan, Yiming Wang, Rui Zhao, Liliang Ren +9 more

This paper proposes Joint Speech-Text Interleaved Pretraining (JSTIP) for speech recognition, which constructs interleaved speech-text sequences and achieves consistent entity accuracy improvement.

View →
cs.SDcs.AIeess.ASRecentMay 29, 2026

Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS

Deokjin Seo, Gangin Park, Kihyun Nam

Chatterbox-Flash introduces a prior-calibrated block diffusion model for zero-shot TTS that achieves high-fidelity, streaming synthesis with significantly lower computational overhead than existing me…

View →
cs.CLcs.AIcs.SDEmpiricalRecentJul 6, 2026

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing

Cheng-Kang Chou, Ming-To Chuang, Ke-Han Lu, Chan-Jan Hsu +1 more

This paper investigates timestamp drift in modern autoregressive ASR systems and proposes REDDIT, a two-stage post-training framework to correct timestamps while avoiding forgetting.

View →
eess.ASEmpiricalRecentJul 20, 2026

X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System

Yuxiang Zhao, Yichi Zhang, Yanjie An, Yanqiao Zhu +9 more

X-Translator is a low-cost modular system for real-time speech-to-speech translation, using streaming ASR, machine translation, and prompt-conditioned TTS, with session-level control.

View →
cs.CLcs.AIcs.SDEmpiricalRecentJun 12, 2026

Learning to Hear Hesitation: Continual Learning for Disfluency-Aware ASR

Henri-Leon Kordt, Theresa Pekarek Rosin, Jae Hee Lee, Stefan Wermter

This paper uses continual learning with explicit disfluency tokens to improve Automatic Speech Recognition (ASR) systems on disfluent speech, addressing the information loss and hallucinations caused…

View →
cs.HCcs.AIcs.SDEmpiricalRecentJun 19, 2026

CORTIS: Text-Only Adaptation of Spoken Language Models for Task-Oriented Voice Agents

Youngwon Choi, Hyeonyu Kim, Taeyoun Kwon, Donghyuk Jung +1 more

CORTIS is a text-only adaptation framework that fine-tunes spoken language models for task-oriented voice agents using text-form task supervision.

View →
eess.AScs.SDEmpiricalRecentAug 7, 2026

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation

Hanke Xie, Haopeng Lin, Jiale Qian, Dake Guo +12 more

This paper proposes SemBridge, a framework for continuous-latent autoregressive speech generation using discrete semantic tokens for supervision.

View →