ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

~ similar to 2606.17662· 15 results

eess.ASEmpiricalRecentJun 19, 2026

Vaani Benchmark V1.0: An Inclusive Multimodal Benchmark Dataset for Hindi

Sujith Pulikodan, Agneedh Basu, Saurabh Kumar, Pranav Bhat +4 more

The paper introduces a new inclusive, multimodal Hindi ASR benchmark with real-world recordings and diverse demographic groups, enabling more robust and realistic evaluation.

View →
cs.CLRecentJun 1, 2026

SN-WER: Script-Normalized WER for Multi-Script Indic ASR Evaluation

Priyaranjan Pattnayak

The paper introduces Script-Normalized WER (SN-WER), a novel evaluation metric that transliterates ASR transcripts into a canonical script to accurately measure speech recognition performance across d…

View →
cs.LGcs.SDEmpiricalRecentJul 24, 2026

Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice Cloning

Roseline Polle, Owen Parsons, George Fairs, Luis Miguel San Martin Fernandez +4 more

Eight voice cloning models are benchmarked on five paralinguistic tasks, showing most preserve signal with modest degradation. Cloning English clinical speech into Japanese outperforms raw cross-lingu…

View →
cs.CLeess.ASEmpiricalRecentJul 3, 2026

Jointly Improving Dialect Identification and ASR in Indian Languages using Multimodal Feature Fusion

Saurabh Kumar, Amartyaveer, Prasanta Kumar Ghosh

This paper proposes a multimodal framework for jointly improving Automatic Speech Recognition (ASR) and Dialect Identification (DID) in Indian languages using a Bottleneck Encoder, RoBERTa encoder, ga…

View →
eess.ASEmpiricalRecentJun 27, 2026

GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

Yujie Tu, Yifan Yang, Tianrui Wang, Yanqiao Zhu +32 more

The paper introduces GigaSpeechBench, a comprehensive multilingual and multidimensional ASR & AST benchmark with 680 hours of human-annotated speech, featuring 12 low-resource languages, 6 Chinese dia…

View →
cs.CLEmpiricalRecentJul 24, 2026

A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books

Varun Ghat Ravikumar, Sina Ahmadi, Lena Jäger, Rico Sennrich

This paper introduces a pipeline to extract grammatical rules, example sentences, and lexicons from grammar books and generates synthetic parallel corpora for fine-tuning machine translation models on…

View →
eess.AScs.AIcs.SDEmpiricalRecentJul 9, 2026

On the Role of Conversational Timing in Synthetic Training Data for ASR

Máté Gedeon, Péter Mihajlik

This paper explores the effect of conversational timing properties on automatic speech recognition (ASR) systems by controlling and optimizing pause and overlap timing distributions.

View →
cs.CLeess.ASEmpiricalRecentJul 2, 2026

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving

Ruchao Fan, Yiming Wang, Rui Zhao, Liliang Ren +9 more

This paper proposes Joint Speech-Text Interleaved Pretraining (JSTIP) for speech recognition, which constructs interleaved speech-text sequences and achieves consistent entity accuracy improvement.

View →
cs.CLcs.CYcs.HCRecentJun 1, 2026

WAXAL-NET: Finetuned Edge ASR Across 19 African Languages

Victor Tolulope Olufemi, Oreoluwa Babatunde, Ramsey Njema, Bolarinwa Gbotemi +27 more

This paper demonstrates that compact, domain-specialized Automatic Speech Recognition (ASR) models significantly outperform large, general-purpose foundation models for conversational speech across 19…

View →
cs.CLcs.SDeess.ASEmpiricalRecentJun 18, 2026

Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal

Syeda Faiza Ahmed Sara, Shammur Absar Chowdhury

A lightweight framework for automated pronunciation assessment using native speech resources and unsupervised or lightly calibrated methods.

View →
cs.SDcs.CLeess.ASEmpiricalRecentJul 19, 2026

Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer

Sivateja Trikutam

This paper presents a method for building a compact Hindi text-to-speech model by pruning a large teacher model under a severe data budget, achieving state-of-the-art performance.

View →
cs.CLcs.AIeess.ASRecentMay 31, 2026

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects

Sicheng Yang, Shulan Ruan, Shiwei Wu, Yu Liu +3 more

PolySpeech-100 introduces a massive, multi-lingual benchmark covering 110 linguistic variants to rigorously test Speech-LLMs, demonstrating that open-source models struggle with low-resource languages…

View →