Publication:

CanTTSo: Improving Cantonese Text-to-Speech with JyutVoice and Zoeng Jyut Gaai Dataset

Loading...
Thumbnail Image

Files

Zhang_William_Thesis.pdf (884.57 KB)

Date

2026-04-16

Journal Title

Journal ISSN

Volume Title

Publisher

Research Projects

Organizational Units

Journal Issue

Access Restrictions

Abstract

Cantonese text-to-speech (TTS) remains challenging and unexplored area within speech production, despite being a language spoken by millions of people worldwide. While recent improvements in neural TTS systems have achieved quality results for high-resource languages, they often lack support for Cantonese or have poor performance because of the tonal complexity and data scarcity.

This thesis presents CanTTSo (Cantonese TTS optimized), a Cantonese text-to-speech system that builds upon the JyutVoice framework through transfer learning from CosyVoice2. Instead of retraining a full end-to-end TTS system from scratch, CanTTSo freezes the 71.3M parameter CosyVoice2 flow matching decoder and trains only a 7.2M parameter Cantonese text encoder and duration predictor on the ZoengjyutGaai storytelling corpus. This approach for Cantonese TTS reduces the memory usage by 3× and the training time by 3–5× compared to full end-to-end training. CanTTSo is evaluated on a 26 sentence test corpus that spans 12 linguistic categories and compared against production-grade TTS systems such as Google Cloud TTS and ElevenLabs Multilingual v2. CanTTSo achieves a Character Error Rate of 0.7914, which outperforms ElevenLabs (0.8063) and approaches Google TTS (0.7425). CanTTSo performs strongest on formal speech and long sentences, which is consistent with the audiobook training domain and generalizes to out-of-domain content with a CER of 0.46. The primary limitation within CanTTSo is the tonal accuracy (0.3653) compared to Google TTS’s 0.6104 and is primarily due to the frozen decoder not encoding the Cantonese tonal distinctions explicitly. Lastly, a speaker similarity score of 0.6096 confirms functional voice cloning capability.

The results illustrate that transfer learning from a large multilingual TTS model is a viable strategy for building Cantonese TTS from a single speaker corpus with greater implications for other low-resource tonal languages such as Taishanese and Hokkien.

Description

Type of resource

Princeton University Senior Theses

Keywords

Location

Citation