ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Familiarity with speech recognition systems”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.CLeess.ASEmpiricalRecentJul 6, 2026

Revisiting the Relation Between Language Model Perplexity and ASR Word Error Rate for Modern End-to-End Speech Recognition

Mohammad Zeineldeen, Albert Zeyer, Haoran Zhang, Robin Schmitt +2 more

This paper investigates the relationship between language model perplexity and word error rate in modern automatic speech recognition systems, studying the impact of external language models, encoder…

View →
eess.AScs.AIcs.SDEmpiricalRecentJul 9, 2026

On the Role of Conversational Timing in Synthetic Training Data for ASR

Máté Gedeon, Péter Mihajlik

This paper explores the effect of conversational timing properties on automatic speech recognition (ASR) systems by controlling and optimizing pause and overlap timing distributions.

View →
eess.AScs.AIcs.SDRecentMay 29, 2026

A Unified and Reproducible Experimentation Framework for Speech Understanding

Jing Peng, Junhao Du, Chenghao Wang, Hanqi Li +20 more

The paper introduces SURE, a unified framework designed to standardize and improve the comparability and reproducibility of evaluations for advanced speech understanding models.

View →
cs.CLcs.SDeess.ASEmpiricalRecentJun 18, 2026

Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal

Syeda Faiza Ahmed Sara, Shammur Absar Chowdhury

A lightweight framework for automated pronunciation assessment using native speech resources and unsupervised or lightly calibrated methods.

View →
eess.ASEmpiricalRecentJun 16, 2026

An Analysis of the Effectiveness of Synthetic Speech Data for ASR Fine-tuning in Selected Indic Languages

Sujith Pulikodan, Agneedh Basu, Pavan Kumar, Pranav Bhat +3 more

This paper investigates the effectiveness of incorporating synthetic speech data in Automatic Speech Recognition (ASR) Systems for three Indic languages by analyzing performance gains, script sources,…

View →
cs.CLRecentMay 28, 2026

Your Multimodal Speech Model Says I Have a Face for Radio

Maya K. Nachesa, Vlad Niculae, Vagrant Gautam

This paper evaluates biases in multimodal speech recognition by testing how pairing different faces with the same audio affects transcription accuracy, finding significant quality-of-service drops acr…

View →
cs.SDcs.AIEmpiricalRecentJun 23, 2026

ZONOS2 Technical Report

Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close +1 more

The authors present ZONOS2 8B, a TTS model with improved naturalness, prosody, and voice cloning fidelity, achieved through scaling, data expansion, and simplification.

View →
cs.SDcs.AIEmpiricalRecentJul 16, 2026

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer +10 more

The paper introduces the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech, speech-to-speech, speech understanding, and automatic speech recognition…

View →
cs.CLRecentJun 1, 2026

SN-WER: Script-Normalized WER for Multi-Script Indic ASR Evaluation

Priyaranjan Pattnayak

The paper introduces Script-Normalized WER (SN-WER), a novel evaluation metric that transliterates ASR transcripts into a canonical script to accurately measure speech recognition performance across d…

View →
eess.ASEmpiricalRecentJun 19, 2026

Vaani Benchmark V1.0: An Inclusive Multimodal Benchmark Dataset for Hindi

Sujith Pulikodan, Agneedh Basu, Saurabh Kumar, Pranav Bhat +4 more

The paper introduces a new inclusive, multimodal Hindi ASR benchmark with real-world recordings and diverse demographic groups, enabling more robust and realistic evaluation.

View →
cs.SDcs.AIEmpiricalRecentJul 21, 2026

What the Waveform Knows: Transparent-first Speech and Audio Intelligence with Caption Studio

Cheng Siong Chin, Jianhua Zhang, Mohan Venkateshkumar

Caption Studio is a transparency-first speech and audio intelligence platform that provides automated transcription, speaker diarization, speech analytics, signal-level audio analysis, and subtitle ge…

View →
cs.LGcs.AIcs.CREmpiricalRecentJun 26, 2026

What Was That Again? Certified Robustness for Automatic Speech Recognition

Andrew C. Cullen, Neil Marchant, Jiani Xie, Paul Montague +1 more

This paper proposes a method to decrease Word Error Rate (WER) and increase recall in Automatic Speech Recognition systems using a certification-inspired mechanism, a Two-Sided Atomic Audit, and a Ran…

View →
eess.ASEmpiricalRecentJun 18, 2026

Transcript-Free Flow-Matching Text-to-Speech via Speech Feature Conditioning

SooHwan Eom, Hee Suk Yoon, Eunseop Yoon, Mark Hasegawa-Johnson +1 more

The paper proposes RTFree-F5, a method to make flow-matching TTS models like F5-TTS independent of reference transcripts, improving performance and naturalness for dysarthric speakers.

View →
cs.CLcs.AIRecentMay 27, 2026

KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs

Haechan Kim, Seungjun Chung, Inkyu Park, Jihoo Lee +1 more

The paper introduces three new Korean speech benchmarks (KVoiceBench, KOpenAudioBench, and KMMAU) to evaluate SpeechLMs, demonstrating that English-centric evaluation fails to capture performance gaps…

View →
cs.SDcs.AIcs.CRRecentJun 4, 2026

Beyond Waveform Robustness: Robust Feature-Vocoder Adversarial Attacks on Automatic Speech Recognition

Yifan Liao, Zongmin Zhang, Zhen Sun, Yuhui Sun +2 more

The paper introduces a novel Clean-Referenced Feature-Vocoder Attack, a black-box adversarial attack that perturbs high-level SSL feature representations instead of raw audio waveforms, achieving supe…

View →
cs.LGcs.SDEmpiricalRecentJul 24, 2026

Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice Cloning

Roseline Polle, Owen Parsons, George Fairs, Luis Miguel San Martin Fernandez +4 more

Eight voice cloning models are benchmarked on five paralinguistic tasks, showing most preserve signal with modest degradation. Cloning English clinical speech into Japanese outperforms raw cross-lingu…

View →