A Practical Guide to Self-Hosting LLMs: Models, Hardware, and Real Cost Math
A software consultant deployed a 70B-class open-weight language model on two GPU servers for a fintech client in January after their compliance team banned all external AI APIs, requiring that no data leave the company's premises. The setup handled document Q&A, support triage, and internal code review at a fraction of the cost of hosted API services. The consultant has since generalized that deployment into a field guide covering model selection, serving stacks, and hardware cost calculations. Open-weight models are categorized into tiers ranging from small 1–4B parameter models requiring 4–8 GB VRAM to large 30–40B models needing 24–48 GB, each suited to different tasks. The guide also cautions that self-hosting is not always cheaper upfront and is best justified by data sovereignty requirements, high inference volumes, or the need for latency predictability.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in