VIDRAFT's VKAE Engine Hits 23x GPU Speedup and 10K Tokens/sec on Nvidia B200
South Korean Pre-AGI startup VIDRAFT has unveiled VKAE, a proprietary inference acceleration engine for large language models. Benchmarked on a single Nvidia B200 GPU, the system achieved roughly 10,000 tokens per second on the Qwen3.5-35B-A3B model, compared to a baseline of approximately 455 tokens per second. This represents up to a 23-fold improvement in GPU utilization without any reported degradation in output quality. VKAE exposes an OpenAI-compatible API, allowing it to integrate seamlessly into existing LLM deployment pipelines. The engine is positioned as competitive with established open-source frameworks like vLLM and TensorRT-LLM, as well as commercial providers such as Groq and Cerebras.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in