An Analysis of the Effectiveness of Synthetic Speech Data for ASR Fine-tuning in Selected Indic Languages
This paper investigates the effectiveness of incorporating synthetic speech data in Automatic Speech Recognition (ASR) Systems for three Indic languages by analyzing performance gains, script sources, speech synthesis models, and voice cloning.
Explores the impact of synthetic speech data on ASR performance using various script sources, speech synthesis models, and voice cloning techniques.
Keywords
Before reading this…
Applications
- →Speech recognition systems
- →Language technology
To understand this paper, make sure you know these concepts first:
- Understanding of Automatic Speech Recognition systemsfind papers →
- Familiarity with speech synthesis and voice cloningfind papers →
Abstract
More Like ThisSynthetic data has the potential to be a valuable resource for training machine learning models, particularly Automatic Speech Recognition (ASR) Systems; however, its effectiveness requires systematic evaluation. In this study, we investigate the impact of incorporating synthetic speech data alongside real-world recordings for three Indic languages: Hindi, Kannada, and Telugu. We analyze the performance gains achieved by augmenting synthetic data with real data and independently examine how ASR performance varies with the sources of scripts used to generate synthetic speech. In addition, we evaluate the effect of synthetic speech generated using different speech synthesis models. Finally, we study the impact of voice cloning in synthetic speech generation on ASR performance, including how performance varies with the number of distinct cloned voices used during data generation.