Q4_K_M Outperforms OpenAI's MXFP4 by 1.8x in Local LLM Speed Test on Apple M2
A developer ran a head-to-head speed test comparing two local LLM quantization formats — Q4_K_M and MXFP4 — on the same Apple M2 MacBook with 24GB unified memory using Ollama. Alibaba's Qwen3-14B in Q4_K_M format averaged 4.7 tokens per second, completing a 200-token task in 44 seconds, while OpenAI's gpt-oss-20B in MXFP4 managed only 2.6 tokens per second, taking nearly 71 seconds for the same workload. Despite MXFP4 being marketed as a faster, low-latency format optimized for local inference, it was approximately 1.8 times slower in warm-run trials. The tester theorizes that the M2 chip lacks native FP4 hardware acceleration — a feature introduced only in the M4 — forcing MXFP4 weights to be dequantized to FP16 before processing, negating its theoretical advantages. The results suggest that MXFP4's performance benefits may be hardware-dependent and not yet fully realized on older Apple Silicon.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in