20 results for “Understanding of self-supervised learning and automatic speech recognition concepts”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
The paper introduces GRIDS, a framework using Local Intrinsic Dimensionality (LID) to detect anomalies in self-supervised speech model representations, showing that LID elevation correlates with ASR d…
This paper examines the absence of phonological context compensation in wav2vec2.0 architecture using Mandarin Chinese tones, contrasting self-supervised pre-training with fine-tuning for ASR.
This paper proposes a novel pretraining method for Acoustic-to-Articulatory Inversion (AAI) using Phoneme Labels, Articulatory Feature Labels, and Critical-articulator Labels, improving performance an…
Xiaoyu Yang, Xuenan Xu, Wenyi Yu, Siyin Wang +9 more
The paper proposes SALMONN-2, an ALLM built on a unified SSL encoder, and presents a multi-layer feature fusion adapter to better exploit hierarchical SSL encoder representations. It also explores mul…
This paper proposes an SSL-AutoEncoder (SSL-AE) approach for reducing feature dimensions in self-supervised learning models while maintaining dysarthric ASR performance.
Yifan Liao, Zongmin Zhang, Zhen Sun, Yuhui Sun +2 more
The paper introduces a novel Clean-Referenced Feature-Vocoder Attack, a black-box adversarial attack that perturbs high-level SSL feature representations instead of raw audio waveforms, achieving supe…
This paper conditions a diffusion-based speech enhancement model on wav2vec 2.0 features using Feature-wise Linear Modulation (FiLM), achieving competitive performance on VoiceBank-DEMAND and LibriMix…
The authors propose a method for phone segmentation and recognition using self-supervised speech models, requiring minimal phonetic transcriptions and generalizing to unseen phones.
A lightweight framework for automated pronunciation assessment using native speech resources and unsupervised or lightly calibrated methods.
The paper introduces Noise-Aware BEATs (NABEATs), a noise-aware audio self-supervised learning framework that estimates clean BEATs representations from noisy audio signals using an auxiliary referenc…
The paper proposes methods for generating global prosodic embeddings using auto-encoder models of pitch and energy, demonstrating competitive or superior performance under challenging conditions.
Mohammad Zeineldeen, Albert Zeyer, Haoran Zhang, Robin Schmitt +2 more
This paper investigates the relationship between language model perplexity and word error rate in modern automatic speech recognition systems, studying the impact of external language models, encoder…
Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close +1 more
The authors present ZONOS2 8B, a TTS model with improved naturalness, prosody, and voice cloning fidelity, achieved through scaling, data expansion, and simplification.
This paper proposes a graph framework using pretrained SSL embeddings for speech analysis in Amyotrophic Lateral Sclerosis (ALS) patients, achieving better results than validation baselines on the SAN…
This paper investigates the correlation between dimensions of self-supervised speech features and speech characteristics, finding that content dimensions primarily capture intensity, formants, and voi…
Jing Peng, Junhao Du, Chenghao Wang, Hanqi Li +20 more
The paper introduces SURE, a unified framework designed to standardize and improve the comparability and reproducibility of evaluations for advanced speech understanding models.
Sujith Pulikodan, Agneedh Basu, Pavan Kumar, Pranav Bhat +3 more
This paper investigates the effectiveness of incorporating synthetic speech data in Automatic Speech Recognition (ASR) Systems for three Indic languages by analyzing performance gains, script sources,…
This paper explores the effect of conversational timing properties on automatic speech recognition (ASR) systems by controlling and optimizing pause and overlap timing distributions.