Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English
This paper fine-tunes two zero-shot text-to-speech models, Chatterbox and CosyVoice 3, on Singlish speakers from the IMDA National Speech Corpus to improve accent similarity and naturalness.
First systematic study of Singlish-accented text-to-speech
Before reading this…
Applications
- →Text-to-speech systems for Singlish language
To understand this paper, make sure you know these concepts first:
- Understanding of zero-shot text-to-speech, Singlish language, and fine-tuning techniquesfind papers →
Abstract
More Like ThisZero-shot text-to-speech (ZS-TTS) achieves near-human quality for standard English, but it copies regional accents poorly. Prompted with a short Singlish utterance, state-of-the-art systems reproduce a speaker's timbre while flattening the accent toward generic English. We investigate whether targeted fine-tuning off-the-shelf ZS-TTS can close the gap for Singapore English (Singlish). We fine-tune two cutting-edge ZS-TTS models, Chatterbox and CosyVoice 3, on 50 Singlish speakers from the IMDA National Speech Corpus. Three speech distributions are evaluated: real recordings against off-the-shelf and fine-tuned generation driven by the same Singlish audio prompts. The evaluation covers four dimensions: naturalness, intelligibility, speaker similarity, and accent similarity. We separate adaptation (in-domain speakers seen during fine-tuning) from consistency (held-out speakers) to test whether accent transfer generalises beyond the training data. Fine-tuning raises accent similarity on in-domain and out-of-domain speakers for both Chatterbox and CosyVoice 3. It moves the generated distribution measurably toward real Singlish, with the gain persisting on held-out speakers. To our knowledge, this is the first systematic study of Singlish-accented TTS.