JaiTTS: Free Thai Voice-Cloning Model Beats Human Benchmark on Short-Text Accuracy
Jasmine Technology Solution (JTS) released JaiTTS-v1.0, a Thai text-to-speech voice-cloning model, publishing their findings on arXiv on April 30, 2026, alongside open-access code and a demo. The model achieved a Character Error Rate (CER) of 1.94% on short-text tasks, narrowly surpassing the human reference score of 1.98% — though it fell short of human performance on longer texts. Built on a modified VoxCPM autoregressive architecture and trained on a large Thai-focused audio corpus, JaiTTS can handle mixed Thai-English text and numerals without requiring text normalization. In a blind listening test involving 20 native Thai evaluators, JaiTTS outperformed two commercial systems across 283 of 400 compared pairs, while losing in 58 pairs. The model also processes audio roughly nine times faster than real time, and is available for free use.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in