Beyond Speaker Independence: Evaluating Cross-Lingual Acoustic-to-Articulatory Inversion Across Finnish and Russian
This paper systematically evaluates acoustic-to-articulatory inversion under domain shifts on FROST-EMA, a Finnish-Russian bilingual EMA corpus, and establishes benchmarks for articulatory targets, acoustic front-ends, and inversion back-ends.
Provides comprehensive benchmarks for acoustic-to-articulatory inversion under domain shifts on a new bilingual corpus
Before reading this…
Applications
- →Speech synthesis
- →Speech recognition
- →Speech therapy
To understand this paper, make sure you know these concepts first:
- Understanding of speech productionfind papers →
- Familiarity with acoustic-to-articulatory inversionfind papers →
Abstract
More Like ThisAcoustic-to-articulatory inversion (AAI) remains challenging under domain shifts where changes in speaker attributes and cross-language conditions often degrade performance. We conduct a systematic evaluation under such shifts and establish baseline benchmarks on FROST-EMA, a Finnish-Russian bilingual EMA corpus. FROST-EMA addresses the English bias and limited speaker diversity of existing resources. We benchmark (i) articulatory targets (raw EMA coordinates vs tract variables), (ii) acoustic front-ends (MFCC vs SSL features), and (iii) inversion back-ends (BiLSTM vs a lightweight attention-based sequence model). We further define evaluation protocols for cross-gender transfer (within language) and cross-language transfer (within gender). The results indicate that cross-gender mismatch introduces moderate Pearson correlation declines (approximately 0.05 to 0.10) relative to the in-domain baseline, whereas cross-language mismatch causes larger drops (approximately 0.10 to 0.20).