27B LLM Squeezed Under 6GB via Ternary Quantization, But CPU Speed Disappoints
A developer tested Prism's Ternary-Bonsai-2-27B, a ternary-quantized large language model that compresses a 27-billion-parameter model into under 6GB, on a budget Hetzner VPS with no GPU. The model uses an aggressive PTQ1_0 format where each weight is reduced to one of three values, requiring Prism's custom llama.cpp fork since the format is not yet supported upstream. While Simon Willison had reported 20–44 tokens per second on Apple Silicon with Metal acceleration, the CPU-only test on the VPS yielded just 2–3 tokens per second. At that speed, the model is viable for background batch processing but impractical for real-time interactive use. The tester also noted an unexplained throughput variance of up to 40 percent between server restarts with no configuration changes.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in