Benchmark of 15 AI API Models Reveals Speed and Cost Gaps Across Providers
A developer published benchmark results comparing 15 AI language model APIs, tested on May 20, 2026, via a unified endpoint across US and Asia-Pacific regions. The tests measured time-to-first-token (TTFT), output speed in tokens per second, and cost per million output tokens using a standardised recursion-explanation prompt. Step-3.5-Flash ranked fastest with a 120ms TTFT and 80 tokens per second at $0.15 per million output tokens, while larger models like Qwen3.5-397B clocked in at 1,200ms TTFT and just 10 tokens per second. The author argues that response latency directly affects user retention, citing a personal case where reducing TTFT improved engagement metrics within two weeks. The findings suggest that several open-weight models offer competitive speed and significantly lower costs compared to proprietary alternatives.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in