Developer finds phone LLM runs 80x slower than theoretical limit
A developer tested Qwen2.5-1.5B-Instruct on a Google Pixel 4 smartphone and measured generation speeds of 0.5 tokens per second. Theoretical limits for the device's Snapdragon 855 processor suggest it should achieve 30-60 tokens per second. The investigation ruled out three common issues: incorrect compiler flags, memory management problems, and API usage errors. The performance gap persists despite optimizations, indicating deeper computational efficiency issues beyond basic configuration.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in