ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

~ similar to 2607.23027· 14 results

cs.CLcs.AIeess.ASRecentMay 31, 2026

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects

Sicheng Yang, Shulan Ruan, Shiwei Wu, Yu Liu +3 more

PolySpeech-100 introduces a massive, multi-lingual benchmark covering 110 linguistic variants to rigorously test Speech-LLMs, demonstrating that open-source models struggle with low-resource languages…

View →
eess.ASEmpiricalRecentJun 27, 2026

GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

Yujie Tu, Yifan Yang, Tianrui Wang, Yanqiao Zhu +32 more

The paper introduces GigaSpeechBench, a comprehensive multilingual and multidimensional ASR & AST benchmark with 680 hours of human-annotated speech, featuring 12 low-resource languages, 6 Chinese dia…

View →
cs.SDcs.AIEmpiricalRecentJun 23, 2026

ZONOS2 Technical Report

Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close +1 more

The authors present ZONOS2 8B, a TTS model with improved naturalness, prosody, and voice cloning fidelity, achieved through scaling, data expansion, and simplification.

View →
cs.CLcs.AIRecentMay 27, 2026

KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs

Haechan Kim, Seungjun Chung, Inkyu Park, Jihoo Lee +1 more

The paper introduces three new Korean speech benchmarks (KVoiceBench, KOpenAudioBench, and KMMAU) to evaluate SpeechLMs, demonstrating that English-centric evaluation fails to capture performance gaps…

View →
eess.AScs.CLcs.LGEmpiricalRecentJun 18, 2026

PASQA: Pitch-Accent-Focused Speech Quality Assessment Model Trained on Synthetic Speech with Accent Errors

Masaya Kawamura, Yuma Shirahata, Kentaro Mitsui, Reo Shimizu

The paper proposes Pitch-Accent-focused Speech Quality Assessment (PASQA) to explicitly target pitch-accent correctness in speech quality assessment, using a controlled Japanese accent-error dataset a…

View →
eess.ASEmpiricalRecentJul 20, 2026

X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System

Yuxiang Zhao, Yichi Zhang, Yanjie An, Yanqiao Zhu +9 more

X-Translator is a low-cost modular system for real-time speech-to-speech translation, using streaming ASR, machine translation, and prompt-conditioned TTS, with session-level control.

View →
cs.CLcs.AIeess.ASEmpiricalRecentJun 16, 2026

Perceptual compensation for tonal context in self-supervised speech models

James Kirby, Ioana Krehan, Michele Gubian

This paper examines the absence of phonological context compensation in wav2vec2.0 architecture using Mandarin Chinese tones, contrasting self-supervised pre-training with fine-tuning for ASR.

View →
eess.AScs.CLcs.LGEmpiricalRecentJun 18, 2026

Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations

Masato Takagi, Masaya Kawamura, Reo Shimizu, Yuma Shirahata

This paper investigates the ability of mean opinion score (MOS) prediction models to capture quality differences in text-to-speech beyond acoustic fidelity through controlled perturbations on speech.

View →
eess.AScs.CLRecentMay 28, 2026

Extracting accent features in spoken Brazilian Portuguese without sociolinguistic labels

Pedro H. L. Leite, Pedro Benevenuto Valadares, Luiz W. P. Biscainho

The paper proposes a novel workflow to extract fine-grained regional accent features in Brazilian Portuguese using only acoustic labels and a phoneme-based forced aligner, showing that localized featu…

View →
cs.CLcs.CYcs.HCRecentJun 1, 2026

WAXAL-NET: Finetuned Edge ASR Across 19 African Languages

Victor Tolulope Olufemi, Oreoluwa Babatunde, Ramsey Njema, Bolarinwa Gbotemi +27 more

This paper demonstrates that compact, domain-specialized Automatic Speech Recognition (ASR) models significantly outperform large, general-purpose foundation models for conversational speech across 19…

View →
eess.AScs.SDEmpiricalRecentJul 28, 2026

Extracting Voice Styles from Frozen TTS Models via Gradient-Based Inverse Optimization

Gyeongmin Kim

The paper describes a method to optimize the style vector for text-to-speech systems without a reference encoder, improving similarity and acceptance rate.

View →