20 results for “Articulatory Feature Labels”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper proposes a novel pretraining method for Acoustic-to-Articulatory Inversion (AAI) using Phoneme Labels, Articulatory Feature Labels, and Critical-articulator Labels, improving performance an…
The study investigates the limitations of Phonetic Foundation Models (PFMs) for Speech-to-IPA transcription using Grapheme-to-Phoneme (G2P) labels and proposes a new approach based on continuous Artic…
The paper proposes a novel workflow to extract fine-grained regional accent features in Brazilian Portuguese using only acoustic labels and a phoneme-based forced aligner, showing that localized featu…
This paper investigates the use of various acoustic features for recognizing dysarthric speech using a Factorized Time Delay Neural Network (F-TDNN) model, achieving a relative improvement of 4.65% in…
A lightweight framework for automated pronunciation assessment using native speech resources and unsupervised or lightly calibrated methods.
SooHwan Eom, Hee Suk Yoon, Eunseop Yoon, Mark Hasegawa-Johnson +1 more
The paper proposes RTFree-F5, a method to make flow-matching TTS models like F5-TTS independent of reference transcripts, improving performance and naturalness for dysarthric speakers.
This paper proposes a multimodal framework for jointly improving Automatic Speech Recognition (ASR) and Dialect Identification (DID) in Indian languages using a Bottleneck Encoder, RoBERTa encoder, ga…
This paper proposes an SSL-AutoEncoder (SSL-AE) approach for reducing feature dimensions in self-supervised learning models while maintaining dysarthric ASR performance.
This paper systematically evaluates acoustic-to-articulatory inversion under domain shifts on FROST-EMA, a Finnish-Russian bilingual EMA corpus, and establishes benchmarks for articulatory targets, ac…
The authors propose a method for phone segmentation and recognition using self-supervised speech models, requiring minimal phonetic transcriptions and generalizing to unseen phones.
Yujie Tu, Yifan Yang, Tianrui Wang, Yanqiao Zhu +32 more
The paper introduces GigaSpeechBench, a comprehensive multilingual and multidimensional ASR & AST benchmark with 680 hours of human-annotated speech, featuring 12 low-resource languages, 6 Chinese dia…
The study investigates the generalization of auto-generated natural-language labels for language model features, finding that while the underlying features show cross-lingual semantic consistency, the…
Eight voice cloning models are benchmarked on five paralinguistic tasks, showing most preserve signal with modest degradation. Cloning English clinical speech into Japanese outperforms raw cross-lingu…
This paper demonstrates that compact, domain-specialized Automatic Speech Recognition (ASR) models significantly outperform large, general-purpose foundation models for conversational speech across 19…
The paper proposes methods for generating global prosodic embeddings using auto-encoder models of pitch and energy, demonstrating competitive or superior performance under challenging conditions.
This paper proposes a multi-level consistency-driven framework, $C^3$ASD, for robust active speaker detection in video, addressing the limitations of recent audio-visual fusion methods.