Maple-Preview runs a 20B ternary MoE AI model at 120 tokens/sec on iPhone
DeepGrove AI has released Maple-Preview, a on-device language model capable of running on an iPhone. The model uses a Mixture of Experts (MoE) architecture with 20 billion parameters and ternary quantization to reduce memory and compute demands. It achieves an inference speed of approximately 120 tokens per second on Apple mobile hardware. The project was shared on Hacker News as a community showcase, drawing early attention from developers interested in on-device AI performance.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in