TensorFold engine delivers 7x speed boost for local LLM inference on Apple Silicon
A developer tested TensorFold, an MIT-licensed inference engine by a contributor named Ash, on a MacBook Pro with an M5 Max chip over a single weekend. The engine achieved up to 220 tokens per second on Qwen3.8-27B, compared to roughly 31 tokens per second with the standard mlx_lm server, using the DFlash2 speculative decoding draft model. Crucially, all drafted outputs matched undrafted outputs byte-for-byte across 16 test cases, confirming the speed gains do not alter model output. Testing revealed the performance advantage stems from TensorFold's own verification mechanism — which checks a full tree of candidate tokens in a single pass — rather than the DFlash2 drafter alone, since llama.cpp with the same drafter reached only a fraction of the speed. The developer also successfully integrated TensorFold behind a Kubernetes service via LLMKube, making the Mac a standard inference node alongside NVIDIA and AMD machines in a mixed cluster.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.



Discussion (0)
Log in to join the discussion and vote.
Log in