Memory Bandwidth, Not VRAM Size, Determines Local LLM Token Speed

When running large language models locally, memory bandwidth in gigabytes per second is the primary factor governing token generation speed, not the total amount of VRAM available. Because autoregressive decoding requires the GPU to stream every model weight from VRAM for each token generated, the theoretical speed ceiling equals memory bandwidth divided by model size. Benchmarks illustrate the stakes sharply: an RTX 4060 Ti running Llama 8B at 42.5 tokens per second dropped to just 3.8 tokens per second when only 20 percent of weights spilled into system RAM, a 91 percent slowdown. Bus width drives this gap — the same 16 GB of VRAM delivers 288 GB/s on a 128-bit RTX 4060 Ti versus 672 GB/s on the 256-bit RTX 4070 Ti Super. For users seeking the best value, a used RTX 3090 with 24 GB on a 384-bit bus at around $650–750 remains the recommended choice for running 32-billion-parameter models with room for context.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in