SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision
This paper introduces a training recipe for sentence-level and long-form streaming speech-to-speech translation using only 2k hours of paired cross-lingual data and auxiliary supervision.
The paper introduces a new training method for long-form streaming speech-to-speech translation using only 2k hours of paired cross-lingual data and auxiliary supervision.
Before reading this…
Applications
- →real-time speech translation
To understand this paper, make sure you know these concepts first:
- understanding of speech-to-speech translationfind papers →
- knowledge of machine learning conceptsfind papers →
Abstract
More Like ThisLong-form streaming speech-to-speech translation (S2ST) requires incremental, unbounded translation under strict latency constraints. Existing methods typically suffer from sentence-bounded supervision or demand massive paired-S2ST supervision. We introduce a training recipe enabling a speech language model for sentence-level and long-form streaming S2ST using only $\sim$2k hours of paired cross-lingual S2ST data, layered atop auxiliary supervision. Anchored by auxiliary multitask training, our approach remains robust even when the paired-S2ST budget itself is reduced by 90\%. Our core contribution, joint text-code trajectory supervision, schedules target text and acoustic semantic codes as a unified commitment path, eliminating the need for separate, unstable speech-side emission controllers. Furthermore, our two-stream Thinker--Talker factorization significantly outperforms unified-decoder baselines by decoupling linguistic reasoning from dense acoustic prediction to mitigate modality interference. Finally, our system achieves highly competitive quality-latency trade-offs on RealSI and ACL60/60-dev, matching state-of-the-art, closed-source S2ST systems such as LiveInterpret~2.0 on ASR-BLEU.