vLLM v0.28.0 brings seven breaking changes that hit small-GPU users hardest
The vLLM inference framework released version 0.28.0 on 26 August 2026, incorporating 584 commits from 270 contributors. Among its most impactful changes for single-GPU users is the removal of built-in bitsandbytes support, which has been moved to an external plugin, meaning anyone using 4-bit or 8-bit quantization must install an additional package or face a broken setup. The default maximum batched token count was also raised from 8,192 to 16,384, which can cause out-of-memory errors on consumer cards with 16GB or 24GB of VRAM. Six additional breaking changes affect users relying on specific Transformers versions, FP8 KV cache scaling, custom attention dtype overrides, reasoning trace parsing, offload metrics dashboards, and legacy MoE kernels. The release also expands CPU and alternative hardware support, enabling DeepSeek-V2/V3 inference on CPU and adding wheels for Intel XPU and IBM s390x platforms.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in