ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “dysarthria”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.SDcs.AIEmpiricalRecentJul 20, 2026

Re-Sonance: A Dysarthric Asynchronous Real-Time Speech Conversion System Based on a Three-Stage Cascaded ASR-LLM-TTS Architecture

Yuxuan Wu, Yifan Xu, Junkun Wang, Jiayong Jiang +2 more

This paper introduces Re-Sonance, a real-time speech-driven AAC system for professional speaking scenarios using LLM-enhanced Whisper ASR, Qwen LLM, and CosyVoice TTS.

View →
eess.AScs.SDNEWEmpiricalJul 28, 2026

Multi-Phonation Graph Learning with Self-Supervised Speech Embeddings for ALS Detection and Progression Prediction

Behrad TaghiBeyglou, Fatemeh Bagheri, Ervin Sejdic

This paper proposes a graph framework using pretrained SSL embeddings for speech analysis in Amyotrophic Lateral Sclerosis (ALS) patients, achieving better results than validation baselines on the SAN…

View →
eess.AScs.AIcs.LGEmpiricalRecentJun 18, 2026

Systematic Study of Dysarthric Speech Recognition: Spectral Features and Acoustic Models

Paban Sapkota, Hemant Kumar Kathania, Mikko Kurimo, Sudarsana Reddy Kadiri +1 more

This paper investigates the use of various acoustic features for recognizing dysarthric speech using a Factorized Time Delay Neural Network (F-TDNN) model, achieving a relative improvement of 4.65% in…

View →
eess.ASEmpiricalRecentJun 23, 2026

Autoencoder based optimized SSL representations: Complexity Minimization and improved Dysarthric ASR

Paban Sapkota, Hemant Kumar Kathania, Mikko Kurimo, Shrikanth Narayanan +1 more

This paper proposes an SSL-AutoEncoder (SSL-AE) approach for reducing feature dimensions in self-supervised learning models while maintaining dysarthric ASR performance.

View →
eess.ASEmpiricalRecentJun 18, 2026

Transcript-Free Flow-Matching Text-to-Speech via Speech Feature Conditioning

SooHwan Eom, Hee Suk Yoon, Eunseop Yoon, Mark Hasegawa-Johnson +1 more

The paper proposes RTFree-F5, a method to make flow-matching TTS models like F5-TTS independent of reference transcripts, improving performance and naturalness for dysarthric speakers.

View →
cs.SDeess.ASEmpiricalRecentJul 26, 2026

Improving Zero-Shot Phonetic Classification through Language-Agnostic Articulatory Features

Ryo Magoshi, Jaeyoung Lee, Shinsuke Sakai, Tatsuya Kawahara

The study investigates the limitations of Phonetic Foundation Models (PFMs) for Speech-to-IPA transcription using Grapheme-to-Phoneme (G2P) labels and proposes a new approach based on continuous Artic…

View →
cs.CLcs.AIcs.SDEmpiricalRecentJun 12, 2026

Learning to Hear Hesitation: Continual Learning for Disfluency-Aware ASR

Henri-Leon Kordt, Theresa Pekarek Rosin, Jae Hee Lee, Stefan Wermter

This paper uses continual learning with explicit disfluency tokens to improve Automatic Speech Recognition (ASR) systems on disfluent speech, addressing the information loss and hallucinations caused…

View →
eess.ASEmpiricalRecentJul 2, 2026

Enhancing Acoustic-to-Articulatory Inversion with Multi-Target Pretraining for Low-Resource Settings

Jesuraj Bandekar, Prasanta Kumar Ghosh

This paper proposes a novel pretraining method for Acoustic-to-Articulatory Inversion (AAI) using Phoneme Labels, Articulatory Feature Labels, and Critical-articulator Labels, improving performance an…

View →
cs.SDcs.AIcs.CRRecentMay 15, 2026

Beyond Content: A Comprehensive Speech Toxicity Dataset and Detection Framework Incorporating Paralinguistic Cues

Zhongjie Ba, Liang Yi, Peng Cheng, Qingcao Li +2 more

The paper introduces ToxiAlert-Bench, a large-scale audio dataset that uniquely annotates both textual and paralinguistic sources of toxicity, and proposes a dual-head neural network that significantl…

View →
cs.LGcs.SDEmpiricalRecentJul 24, 2026

Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice Cloning

Roseline Polle, Owen Parsons, George Fairs, Luis Miguel San Martin Fernandez +4 more

Eight voice cloning models are benchmarked on five paralinguistic tasks, showing most preserve signal with modest degradation. Cloning English clinical speech into Japanese outperforms raw cross-lingu…

View →
eess.ASEmpiricalRecentJun 27, 2026

GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

Yujie Tu, Yifan Yang, Tianrui Wang, Yanqiao Zhu +32 more

The paper introduces GigaSpeechBench, a comprehensive multilingual and multidimensional ASR & AST benchmark with 680 hours of human-annotated speech, featuring 12 low-resource languages, 6 Chinese dia…

View →
eess.AScs.CLcs.LGEmpiricalRecentJun 18, 2026

Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations

Masato Takagi, Masaya Kawamura, Reo Shimizu, Yuma Shirahata

This paper investigates the ability of mean opinion score (MOS) prediction models to capture quality differences in text-to-speech beyond acoustic fidelity through controlled perturbations on speech.

View →
cs.SDcs.CLEmpiricalRecentJun 18, 2026

Segment-Level Mandarin Chinese Speech-Based Cognitive Impairment Detection via an Autoencoder with Contrastive Learning

Yongqi Shao, Hong Huo, Flavio Bertini, Danilo Montesi +1 more

This paper proposes a segment-level representation learning framework for speech-based cognitive impairment detection using autoencoders and contrastive objectives.

View →
cs.CLcs.SDeess.ASEmpiricalRecentJun 18, 2026

Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal

Syeda Faiza Ahmed Sara, Shammur Absar Chowdhury

A lightweight framework for automated pronunciation assessment using native speech resources and unsupervised or lightly calibrated methods.

View →
cs.CLeess.ASEmpiricalRecentJul 6, 2026

Revisiting the Relation Between Language Model Perplexity and ASR Word Error Rate for Modern End-to-End Speech Recognition

Mohammad Zeineldeen, Albert Zeyer, Haoran Zhang, Robin Schmitt +2 more

This paper investigates the relationship between language model perplexity and word error rate in modern automatic speech recognition systems, studying the impact of external language models, encoder…

View →