LLMs Now Run Locally on $8 Microcontrollers at Near-Reading Speed
On July 25, 2026, a developer known as slvDev published esp32-ai, demonstrating a 28.9-million-parameter language model running entirely on an ESP32-S3 microcontroller at 9.88 tokens per second — roughly human reading speed. The feat was made possible not by new hardware but by a memory-hierarchy architecture inspired by Google's Gemma 3n, splitting the model so a small reasoning core sits in RAM while a large embedding table is accessed in tiny slices from flash storage. The project drew coverage from Tom's Hardware, The Register, and others, representing a hundredfold jump from the previous ESP32 parameter record of 260,000. A lesser-noticed follow-up project, p-for-llm, went further on August 5, running a 180.9-million-parameter mixture-of-experts model on the similarly priced ESP32-P4 at around 9 tokens per second, with basic instruction-following and an early tool-calling capability. Together, the two projects mark a significant shift in what edge AI hardware can achieve without any cloud offloading.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in