Autoencoder based optimized SSL representations: Complexity Minimization and improved Dysarthric ASR
This paper proposes an SSL-AutoEncoder (SSL-AE) approach for reducing feature dimensions in self-supervised learning models while maintaining dysarthric ASR performance.
Introducing SSL-AE bottlenecking as an effective solution for SSL feature compression in resource-constrained environments
Keywords
Before reading this…
Applications
- →Dysarthric Automatic Speech Recognition (ASR)
To understand this paper, make sure you know these concepts first:
- Understanding of self-supervised learning and automatic speech recognition conceptsfind papers →
Abstract
More Like ThisSelf-supervised learning (SSL) models extract rich speech representations but often come with high-dimensional features, increasing computational complexity. This work explores an SSL-AutoEncoder (SSL-AE) bottlenecking approach to efficiently reduce feature dimensions while maintaining dysarthric Automatic Speech Recognition (ASR) performance. By leveraging an autoencoder, we transform high-dimensional SSL features into a compact space, reducing model complexity and training time. Our method preserves essential speech information, achieving reduced Word Error Rates (WER) while significantly lowering computational costs. Experiments show SSL-AE bottlenecking reduces training time by 8x compared to the SSL baseline, demonstrating efficiency without sacrificing recognition performance. These results highlight AE as an effective solution for SSL feature compression in resource-constrained environments.