ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Understanding of text-to-speech models”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

eess.AScs.AIcs.SDRecentMay 29, 2026

A Unified and Reproducible Experimentation Framework for Speech Understanding

Jing Peng, Junhao Du, Chenghao Wang, Hanqi Li +20 more

The paper introduces SURE, a unified framework designed to standardize and improve the comparability and reproducibility of evaluations for advanced speech understanding models.

View →
cs.CLcs.LGcs.SDEmpiricalRecentJun 21, 2026

Interleaved Speech Language Models Latently Work In Text

Talia Sternberg, Gallil Maimon, Yossi Adi

This paper analyzes speech-text interleaved language models and reveals that they go through an implicit transcription phase in which spoken words become decodable as text in intermediate layers.

View →
cs.CLeess.ASEmpiricalRecentJul 6, 2026

Revisiting the Relation Between Language Model Perplexity and ASR Word Error Rate for Modern End-to-End Speech Recognition

Mohammad Zeineldeen, Albert Zeyer, Haoran Zhang, Robin Schmitt +2 more

This paper investigates the relationship between language model perplexity and word error rate in modern automatic speech recognition systems, studying the impact of external language models, encoder…

View →
cs.CLcs.AIeess.ASRecentMay 31, 2026

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects

Sicheng Yang, Shulan Ruan, Shiwei Wu, Yu Liu +3 more

PolySpeech-100 introduces a massive, multi-lingual benchmark covering 110 linguistic variants to rigorously test Speech-LLMs, demonstrating that open-source models struggle with low-resource languages…

View →
eess.AScs.CLcs.LGEmpiricalRecentJun 18, 2026

Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations

Masato Takagi, Masaya Kawamura, Reo Shimizu, Yuma Shirahata

This paper investigates the ability of mean opinion score (MOS) prediction models to capture quality differences in text-to-speech beyond acoustic fidelity through controlled perturbations on speech.

View →
cs.SDcs.AIEmpiricalRecentJun 23, 2026

ZONOS2 Technical Report

Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close +1 more

The authors present ZONOS2 8B, a TTS model with improved naturalness, prosody, and voice cloning fidelity, achieved through scaling, data expansion, and simplification.

View →
cs.SDcs.AIRecentJun 1, 2026

MOSS-Audio Technical Report

Chen Yang, Chufan Yu, Hanfu Chen, Jie Zhu +21 more

MOSS-Audio is a unified audio-language model designed for comprehensive understanding of speech, environmental sounds, and music, achieving strong performance across various audio-grounded tasks.

View →
cs.CLcs.AIRecentMay 27, 2026

KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs

Haechan Kim, Seungjun Chung, Inkyu Park, Jihoo Lee +1 more

The paper introduces three new Korean speech benchmarks (KVoiceBench, KOpenAudioBench, and KMMAU) to evaluate SpeechLMs, demonstrating that English-centric evaluation fails to capture performance gaps…

View →
eess.ASEmpiricalRecentJun 18, 2026

Transcript-Free Flow-Matching Text-to-Speech via Speech Feature Conditioning

SooHwan Eom, Hee Suk Yoon, Eunseop Yoon, Mark Hasegawa-Johnson +1 more

The paper proposes RTFree-F5, a method to make flow-matching TTS models like F5-TTS independent of reference transcripts, improving performance and naturalness for dysarthric speakers.

View →
cs.HCcs.AIcs.SDEmpiricalRecentJun 19, 2026

CORTIS: Text-Only Adaptation of Spoken Language Models for Task-Oriented Voice Agents

Youngwon Choi, Hyeonyu Kim, Taeyoun Kwon, Donghyuk Jung +1 more

CORTIS is a text-only adaptation framework that fine-tunes spoken language models for task-oriented voice agents using text-form task supervision.

View →
eess.ASEmpiricalRecentJun 16, 2026

An Analysis of the Effectiveness of Synthetic Speech Data for ASR Fine-tuning in Selected Indic Languages

Sujith Pulikodan, Agneedh Basu, Pavan Kumar, Pranav Bhat +3 more

This paper investigates the effectiveness of incorporating synthetic speech data in Automatic Speech Recognition (ASR) Systems for three Indic languages by analyzing performance gains, script sources,…

View →
eess.AScs.SDEmpiricalRecentJul 28, 2026

Extracting Voice Styles from Frozen TTS Models via Gradient-Based Inverse Optimization

Gyeongmin Kim

The paper describes a method to optimize the style vector for text-to-speech systems without a reference encoder, improving similarity and acceptance rate.

View →
cs.SDcs.AIeess.ASRecentMay 29, 2026

Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS

Deokjin Seo, Gangin Park, Kihyun Nam

Chatterbox-Flash introduces a prior-calibrated block diffusion model for zero-shot TTS that achieves high-fidelity, streaming synthesis with significantly lower computational overhead than existing me…

View →
eess.AScs.SDEmpiricalRecentJul 7, 2026

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs

Ke-Han Lu, Keqi Deng, Ruchao Fan, Rui Zhao +1 more

The paper proposes SpeechKV, a method to compress speech sequences inside large language models using a learned pooling, maintaining performance and delivering decoding speedup.

View →
cs.SDcs.AIeess.ASRecentMay 28, 2026

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

Bohan Li, Shi Lian, Hankun Wang, Yiwei Guo +5 more

HoliTok introduces a novel continuous holistic tokenization model that provides a unified, high-fidelity latent representation for simultaneously supporting both speech generation and speech understan…

View →
eess.AScs.AIcs.SDEmpiricalRecentJul 9, 2026

On the Role of Conversational Timing in Synthetic Training Data for ASR

Máté Gedeon, Péter Mihajlik

This paper explores the effect of conversational timing properties on automatic speech recognition (ASR) systems by controlling and optimizing pause and overlap timing distributions.

View →