Developer Publishes Hands-On Guide to Running 594B-Parameter LLMs on Local Hardware
Software developer James O'Beirne has published a detailed GitHub repository documenting his real-world experience building and configuring local hardware to run state-of-the-art large language models. The guide covers two budget tiers: a roughly $2,000 setup using two used RTX 3090 GPUs capable of running Qwen3-27B, and a $40,000 configuration built around four RTX PRO 6000 Blackwell cards with 384GB of combined VRAM. To keep costs down on the high-end build, O'Beirne sourced most components from eBay, including a last-generation EPYC Milan server board, bringing the base system cost to around $5,587. The guide goes beyond hardware specs, detailing critical BIOS settings, kernel flags, and PCIe switch configurations needed to enable proper multi-GPU communication, including fixes for issues like ASPM throttling and ACS silently rerouting peer-to-peer traffic through the CPU. O'Beirne also shares practical tooling setups covering model serving via Docker Compose, local speech-to-text using Whisper, and a sandboxed VM environment for running AI coding agents.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in