Chinese AI Models Take Up to 34 Seconds Per Response, Far Behind Western Rivals
A developer benchmarking AI models for a text-based game found that Chinese large language models are significantly slower than their Western counterparts, with response times ranging from 14 to 34 seconds for a standard prompt. Models tested included Kimi K3, MiniMax M3, Qwen variants, and DeepSeek V4, all evaluated on the same 36,000-character prompt representing 8,000–13,000 tokens. In comparison, leading Western models such as Claude 5 Opus, GPT-5.6, and Mistral Large responded in roughly 3 to 8 seconds using the same input. The developer also noted that Chinese models consistently generated far more output tokens, and that Qwen models required capped reasoning limits just to stay under 30 seconds. A side finding revealed that Claude's tokenizer billed 63% more prompt tokens than Google's for identical text, highlighting significant cost differences across providers.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in