~ similar to 2607.26751· 14 results
This paper analyzes speech-text interleaved language models and reveals that they go through an implicit transcription phase in which spoken words become decodable as text in intermediate layers.
The authors propose a method for phone segmentation and recognition using self-supervised speech models, requiring minimal phonetic transcriptions and generalizing to unseen phones.
Youngwon Choi, Hyeonyu Kim, Taeyoun Kwon, Donghyuk Jung +1 more
CORTIS is a text-only adaptation framework that fine-tunes spoken language models for task-oriented voice agents using text-form task supervision.
Sicheng Yang, Shulan Ruan, Shiwei Wu, Yu Liu +3 more
PolySpeech-100 introduces a massive, multi-lingual benchmark covering 110 linguistic variants to rigorously test Speech-LLMs, demonstrating that open-source models struggle with low-resource languages…
The paper demonstrates that the location and nature of state encoding in sequence models are not fixed architectural traits but are highly dependent on the specific task, showing that the encoding pro…
Mohammad Zeineldeen, Albert Zeyer, Haoran Zhang, Robin Schmitt +2 more
This paper investigates the relationship between language model perplexity and word error rate in modern automatic speech recognition systems, studying the impact of external language models, encoder…
The paper proposes DOA, a training-free attention policy that leverages self-attention in decoder-only SpeechLLMs to achieve high-quality, low-latency simultaneous long-form translation without requir…
The paper quantitatively confirms the Currier A/B language distinction in the Voynich Manuscript, demonstrating it is governed by a higher-dimensional, context-dependent boolean switch rather than a s…
This paper systematically evaluates acoustic-to-articulatory inversion under domain shifts on FROST-EMA, a Finnish-Russian bilingual EMA corpus, and establishes benchmarks for articulatory targets, ac…
This paper introduces Voice Memory, an inference-only scheme for agentic speech recognition that uses a frozen corrector and a score-gated optimizer to improve speech recognition performance.
A lightweight framework for automated pronunciation assessment using native speech resources and unsupervised or lightly calibrated methods.
The paper demonstrates that an attention-augmented LSTM model can achieve near-perfect character-level decipherment of homophonic ciphertexts from historical English and Swedish, even under challengin…
Yushan Yashengjiang, Jie Zhang, Miao Sun, Huadong Liang +2 more
This paper proposes an end-to-end Markov framework for auditory attention decoding using conditional random fields and an EEG--speech correlation backbone.