Qwen3-8B Benchmark: vLLM Leads SGLang and llama.cpp on Workstation Blackwell GPU
A hands-on benchmark tested the Qwen3-8B language model across three inference stacks — vLLM 0.27.1, SGLang 0.5.9, and llama.cpp — on an RTX PRO 6000 Blackwell workstation GPU with 96 GB VRAM. At concurrency 32, vLLM delivered the highest throughput at 1,725 tokens per second, roughly 1.3x faster than SGLang and 4x faster than llama.cpp, with time-to-first-token also 4x lower than llama.cpp. For single-user workloads, however, all three stacks performed similarly in the 83–96 tokens-per-second range, making stack choice largely irrelevant at low concurrency. Switching vLLM to an FP8 official checkpoint yielded a clean 1.5x throughput gain — rising to 2,597 tokens per second aggregate — with no factual regressions detected across a 20-prompt quality check. The author also flagged that workstation Blackwell hardware requires workarounds for kernel-level issues not present on datacenter GPUs, advising users on RTX PRO 6000 or consumer Blackwell to budget extra setup time.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in