Magnitude launches self-optimizing local AI inference engine, claims 2x llama.cpp speed
YC S25 startup Magnitude, founded by engineers Anders and Tom, has launched an open-source inference engine designed specifically for running AI agents on local hardware. Unlike existing solutions such as llama.cpp or vLLM, Magnitude uses on-device kernel compilation and tuning to maximize performance on the user's specific hardware across Mac, Linux, and Windows. The engine employs dynamic memory allocation and hybrid paged attention to support multiple concurrent agent sessions without monopolizing system resources. Benchmarks against llama.cpp using Qwen 3.6 35B show up to 92% faster decode speeds on Apple M4 Pro and 23% faster prefill on CUDA hardware, alongside roughly 27-28% lower per-agent memory usage. Built in Rust and licensed under Apache 2.0, Magnitude ships as a desktop app and integrates with popular agent tools, with future plans including loading oversized models from RAM or disk just-in-time.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in