Thai researchers release free local-run Thai TTS model with zero-shot voice cloning
ThonburianTTS is an open-source Thai text-to-speech model developed by researchers at Mahidol University's Biomedical and Data Lab in collaboration with Looloo Technology. Built on the F5-TTS flow-matching architecture and fine-tuned on the GigaSpeech2 Thai dataset, it can run entirely on a local machine without requiring any external API. The model supports zero-shot voice cloning using just 5–10 seconds of sample audio, achieving similarity scores of 84–89% according to the original research paper. It was presented at the iSAI-NLP 2025 conference and is freely available on HuggingFace and GitHub under a MIT code licence, though the model weights carry a CC BY-NC-SA 4.0 licence that prohibits commercial use. Researchers note the model struggles with longer text passages, a limitation attributed to its training data favouring short utterances.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in