How to Run Qwen 3 27B Locally on an RTX 3090 Using Unsloth and DeepSeek Harness
Developer Jacques Gariépy has published a detailed technical guide on running the Qwen 3 27B language model entirely on a local Windows 11 machine using an NVIDIA RTX 3090 GPU with 24 GB of VRAM. The setup combines three components: the DeepSeek Harness open-source agent framework, the Unsloth inference engine built on llama.cpp with CUDA 13 support, and the Qwen 3 27B model in UD-Q4_K_XL dynamic quantization format. With this configuration, the model occupies approximately 17.5 GB of VRAM for weights and 5 GB for a 32,000-token KV cache, keeping total usage near 23.5 GB and avoiding slower system RAM. The guide covers step-by-step installation, common Windows-specific bugs such as the SSLKEYLOGFILE issue, server startup, and usage via both a web interface and a CLI mode styled after Claude Code. The approach enables fully private, zero-API-cost AI inference with minimal latency on consumer hardware.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in