ArXivCSExplorer
β˜†β˜†BookmarksπŸ†RSSHow to UseFAQ
Built with and by Teycir Ben Soltaneβ€’
How to Useβ€’FAQβ€’GitHubβ€’arXiv.orgβ€’
Share:

20 results for β€œTTS”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.β“˜

Want pure semantic search? Try claim verification β†’

cs.SDEmpiricalRecentJul 22, 2026

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis

Kaicheng Luo, Xuefei Gong, Yutao Sun, Jinling He +5 more

This paper introduces StellarTTS, a mobile-optimized non-autoregressive text-to-speech framework with sparse temporal embeddings and a semantic-aware codec, achieving lower latency and stronger robust…

View β†’
eess.ASEmpiricalRecentJun 18, 2026

Transcript-Free Flow-Matching Text-to-Speech via Speech Feature Conditioning

SooHwan Eom, Hee Suk Yoon, Eunseop Yoon, Mark Hasegawa-Johnson +1 more

The paper proposes RTFree-F5, a method to make flow-matching TTS models like F5-TTS independent of reference transcripts, improving performance and naturalness for dysarthric speakers.

View β†’
eess.ASEmpiricalRecentJul 25, 2026

Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English

Ivan Kukanov, Zheng Xin Chai

This paper fine-tunes two zero-shot text-to-speech models, Chatterbox and CosyVoice 3, on Singlish speakers from the IMDA National Speech Corpus to improve accent similarity and naturalness.

View β†’
cs.SDcs.AIEmpiricalRecentJun 23, 2026

ZONOS2 Technical Report

Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close +1 more

The authors present ZONOS2 8B, a TTS model with improved naturalness, prosody, and voice cloning fidelity, achieved through scaling, data expansion, and simplification.

View β†’
cs.SDcs.AIeess.ASRecentMay 29, 2026

Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS

Deokjin Seo, Gangin Park, Kihyun Nam

Chatterbox-Flash introduces a prior-calibrated block diffusion model for zero-shot TTS that achieves high-fidelity, streaming synthesis with significantly lower computational overhead than existing me…

View β†’
eess.AScs.AIcs.CLRecentMay 29, 2026

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment

Jun-Hak Yun, Seung-Bin Kim, Seong-Whan Lee

ImmersiveTTS is an environment-aware text-to-speech model that generates natural speech seamlessly integrated within environmental contexts by explicitly modeling cross-modal interactions, achieving s…

View β†’
eess.ASDatasetRecentJul 15, 2026

Dialogs: a studio-quality expressive conversational Russian speech corpus for dialog assistants

Ilya Shigabeev, Ilya Latyshev

The paper introduces Dialogs, a new Russian conversational speech corpus with high-quality recordings, segmented utterances, and expressive prosody labels.

View β†’
cs.SDcs.AIEmpiricalRecentJul 20, 2026

Re-Sonance: A Dysarthric Asynchronous Real-Time Speech Conversion System Based on a Three-Stage Cascaded ASR-LLM-TTS Architecture

Yuxuan Wu, Yifan Xu, Junkun Wang, Jiayong Jiang +2 more

This paper introduces Re-Sonance, a real-time speech-driven AAC system for professional speaking scenarios using LLM-enhanced Whisper ASR, Qwen LLM, and CosyVoice TTS.

View β†’
eess.ASEmpiricalRecentJul 20, 2026

X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System

Yuxiang Zhao, Yichi Zhang, Yanjie An, Yanqiao Zhu +9 more

X-Translator is a low-cost modular system for real-time speech-to-speech translation, using streaming ASR, machine translation, and prompt-conditioned TTS, with session-level control.

View β†’
eess.AScs.SDEmpiricalRecentJul 28, 2026

Extracting Voice Styles from Frozen TTS Models via Gradient-Based Inverse Optimization

Gyeongmin Kim

The paper describes a method to optimize the style vector for text-to-speech systems without a reference encoder, improving similarity and acceptance rate.

View β†’
eess.ASEmpiricalRecentJul 20, 2026

The tttAI System for the TSA-ASR Task of the SmartGlasses Challenge 2026

Xuanji He, Gaoyang Dong, Xiaoxiao Li, Minchuan Chen +1 more

The paper presents the tttAI system for time-stamped speaker-attributed speech recognition in smart-glasses recordings, achieving a tcpCER of 7.10% on Track 1 and 34.04% on Track 2.

View β†’
eess.AScs.AIRecentMay 29, 2026

OpenSTBench: Beyond Semantic Evaluation for Speech Translation

Yanjie An, Yuxiang Zhao, Yichi Zhang, Qixi Zheng +4 more

The paper introduces OpenSTBench, a unified, multidimensional evaluation framework designed to comprehensively compare heterogeneous speech translation systems by jointly assessing translation, speech…

View β†’
cs.CLeess.ASEmpiricalRecentJul 2, 2026

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving

Ruchao Fan, Yiming Wang, Rui Zhao, Liliang Ren +9 more

This paper proposes Joint Speech-Text Interleaved Pretraining (JSTIP) for speech recognition, which constructs interleaved speech-text sequences and achieves consistent entity accuracy improvement.

View β†’
eess.AScs.SDEmpiricalRecentJul 16, 2026

SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

Shuai Wang, Zihan Qian, Ke Zhang, Jiangyu Han +8 more

This paper introduces the REAL-TSE Challenge, a satellite challenge on target speaker extraction from real conversational recordings, and describes its task definition, datasets, evaluation protocol,…

View β†’
cs.SDEmpiricalRecentJul 22, 2026

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision

Rongshen He, Xinyu Liang, Dekun Chen, Jiaqi Li +2 more

This paper introduces a training recipe for sentence-level and long-form streaming speech-to-speech translation using only 2k hours of paired cross-lingual data and auxiliary supervision.

View β†’
eess.ASEmpiricalRecentJun 27, 2026

GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

Yujie Tu, Yifan Yang, Tianrui Wang, Yanqiao Zhu +32 more

The paper introduces GigaSpeechBench, a comprehensive multilingual and multidimensional ASR & AST benchmark with 680 hours of human-annotated speech, featuring 12 low-resource languages, 6 Chinese dia…

View β†’