Developers Build 100MB Voice-Cloning TTS Model That Runs Offline on Budget Phones
A team has developed a text-to-speech model with voice cloning capability that weighs under 100MB and runs entirely on-device, requiring no cloud or API connection at inference time. The model uses roughly 45 million parameters, combining a language model and an audio decoder, making it far smaller than typical server-side TTS systems that can reach several gigabytes. Early tests on a budget Android phone achieved faster-than-realtime speech synthesis on a single CPU core, with full benchmarks across multiple devices planned for a future post. Voice cloning requires only a five-second audio sample, with the cloning step handled off-device while all speech synthesis happens locally, even on hardware as modest as a Raspberry Pi. The team overcame key training challenges, including a tendency for single-step flow samplers to produce flat, monotone output and instability from mixed-precision training, both of which required custom solutions to resolve.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in