Qwen3 27B Runs on One RTX 3090; Dual-Card Setup Offers Modest Speed Gains

The Qwen3 27.8B dense model, released on August 14, can run on a single RTX 3090 GPU using W4A16 quantization, achieving decode speeds of roughly 125–155 tokens per second. A home-lab benchmark compared single-card and dual-card (tensor parallel) configurations using two RTX 3090s running vLLM. Adding a second card yielded approximately 20% faster decoding and 14% faster cold prefill on average, with larger prompts seeing up to 23% improvement in time-to-first-token. Cached prompts appeared 2–3x faster on both setups, but this was attributed to vLLM's prefix caching rather than the extra GPU. At current API pricing and measured electricity costs, generating one million output tokens costs the equivalent of roughly $0.12–$0.20 in power on the test machine.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in