Nari Labs launches fast, cheap Qwen3 TTS and ASR endpoints beating closed rivals
Nari Labs, founded by Toby and known for the open-source Dia text-to-speech model, has released optimized inference endpoints for Alibaba's Qwen3-TTS and Qwen3-ASR speech models. The company built a specialized inference engine that runs Qwen3-TTS at under 50ms latency at 10 requests per second, outperforming general-purpose systems like vLLM and SGLang. According to Coval's voice AI benchmarks, Nari's Qwen3-TTS endpoint ranks first in word error rate accuracy and second in latency among providers including ElevenLabs and Cartesia, while also being the cheapest option. Their Qwen3-ASR endpoint claims the lowest latency and second-best accuracy, priced as the second cheapest on the benchmark list. Nari Labs has open-sourced the TTS inference engine on GitHub and says it plans to expand into audio diarization, video, and world model inference.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in