Cactus releases Needle 2, a 14MB AI model built for phones, wearables, and IoT devices
Cactus has launched Needle 2, a 45-million-parameter language model compressed to just 14MB that runs on resource-constrained devices including budget smartphones, Raspberry Pis, wearables, and small robots. The model operates within 28MB of RAM and achieves 300–500 tokens per second on sub-$200 Android phones and single-board computers, targeting the roughly 21 billion connected IoT devices worldwide that lack dedicated AI hardware. Needle 2 is designed specifically for structured tasks such as tool calling, device control, and data extraction, with the developers arguing that these functions require no broad world knowledge, making a small parameter count sufficient. On relevant benchmarks, it competes with models five to seventy times larger, including Apple's Foundation Model and LFM2.5 230M, while consuming significantly fewer compute resources per token. The model supports fine-tuning on a standard laptop in minutes to a few hours and includes a confidence-scoring mechanism that can escalate low-certainty requests to a larger cloud-based model.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in