Prism ML Bonsai 2 27B Squeezes Qwen3 into 5.9GB, but Benchmarks Have Caveats

Prism ML released Bonsai 2 27B, a ternary-quantized version of the Qwen3 27B model compressed to just 5.9GB — roughly one-ninth of the original size — using 1.76-bit-per-weight encoding. The company claims the model retains 98.2% of the original's aggregate benchmark score, with math and coding performance largely preserved, though visual reasoning showed the steepest drop. Independent analyst Kaitchup raised concerns that the 98.2% figure is based on average scores and may not reflect real-world multi-step tasks like agentic coding, where small errors compound. Benchmarks were also run on a full-precision model via vLLM on H100 GPUs, not on the compressed GGUF files users would actually run locally. On the hardware side, the smaller file size translates to roughly 9x faster token generation, with an RTX 5090 outpacing even the H100 at 129.9 versus 113.9 tokens per second at batch size 1.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in