Laptop GPU Outperforms 12-Core CPU by 4.3x Running Google Gemma 4 Model
A developer benchmarked Google's Gemma 4 E2B language model on a single laptop equipped with a 13th-gen Intel Core i7-1360P and an Nvidia GTX 1650 Ti (4 GB) GPU. The test used the llama.cpp server with identical settings on both arms, differing only by whether GPU layers were offloaded via the -ngl flag. Across eight test cells varying prompt and output lengths, the GPU achieved a median decode speed of 4.27 times faster than the CPU, with individual ratios ranging from 4.04x to 4.34x. The GPU also delivered prefill times roughly 3.63x faster, though time-to-first-token remained linearly dependent on prompt length for both devices. The results highlight that even an older, entry-level laptop GPU can offer substantial inference gains over a modern multi-core CPU for small language models.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in