AI PC Label Misleads: NPU Skips LLM Inference, Leaving CPU to Struggle
A hands-on benchmark of three consumer machines running Gemma 4 26B (18 GB) reveals that LLM inference splits into two phases — prefill and generation — with very different hardware demands. Prefill, which processes the input prompt before any output is generated, is compute-bound and varies nearly 18-fold across tested machines, while generation speed differs by less than twofold. A laptop marketed as an 'AI PC,' powered by AMD's Ryzen 8840U with a dedicated XDNA NPU, failed at large-prompt tasks because current LLM runners do not utilise the NPU, forcing the CPU to handle prefill at just 20 tokens per second and causing stalls exceeding 13 minutes on long prompts. Separately, model-load times were found to be storage-bound, with a machine using a SATA SSD taking over 50 seconds to load the 18 GB model versus around 8 seconds on NVMe, making a long keep-alive setting essential on slower-disk systems. The findings suggest that advertised AI PC specifications, including TOPS ratings, are largely irrelevant for running large language models against long prompts.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in