Qwen 3.8 27B 1-bit Quantization Delivers Under 1 Token/sec on RTX 4090
A developer benchmarked Qwen 3.8 27B in both 4-bit and 1-bit quantization formats on an NVIDIA RTX 4090 and a MacBook Pro M1 Max to evaluate real-world inference speeds for local AI agent use cases. The 4-bit (q4_K_M) model consumed around 9.5GB of VRAM and averaged 28.5 tokens per second, making it practical for agent workflows requiring fast responses. By contrast, the 1-bit (q1_K) model used only 4.5GB of VRAM but averaged a mere 0.8 tokens per second — effectively unusable for any interactive or automated task. Testing was conducted using Ollama 0.1.37, with each configuration run 10 times on a standardized coding prompt to ensure consistent results. The findings challenge the assumption that lower VRAM usage translates to faster inference, highlighting that aggressive quantization can severely degrade throughput despite memory savings.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in