Laptop GPU Outpaces Its Own CPU by 4.1x Running Google Gemma 4 AI Model
A developer benchmarked a GTX 1650 Ti GPU against an Intel Core i7-10750H CPU on the same Lenovo Yoga 9 laptop, both serving Google's Gemma 4 E2B model via llama.cpp on Debian sid. The GPU decoded tokens 4.14 times faster, prefilled 3.42 times faster, and completed requests 3.62 times end-to-end faster than the CPU. To eliminate thermal bias from the shared cooling system, tests were run in ABBA order — alternating which device went first — revealing that single-order benchmarks on this machine carry roughly a 2% error. Both builds used the same llama.cpp commit and identical quantized model file, differing only by a single flag enabling or disabling CUDA. The author also released a suite of Python MCP tools on GitHub to help manage llama.cpp server deployments.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in