Qwen3 27B Model Runs at 50 Tokens/s with 256K Context on a 24GB GPU
A developer has demonstrated running the Qwen3 27B large language model on a single 24GB NVIDIA RTX PRO 4000 SFF graphics card. The setup achieves a throughput of approximately 50 tokens per second using Multi-Token Prediction (MTP). Notably, the configuration supports an extended 256K token context window, which is unusually large for consumer-grade hardware. The RTX PRO 4000 SFF offers 432 GB/s of memory bandwidth, which appears to be a key factor enabling this performance. The findings suggest that capable large models can be run efficiently on relatively accessible hardware given the right configuration.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in