ExecuTorch MLX Delegate Runs Qwen3 Up to 4.52x Faster Than PyTorch MPS on Apple Silicon
A developer tested Apple Silicon inference speeds using ExecuTorch 1.3.1's experimental MLX delegate, released May 18, 2026, with the Qwen3-0.6B language model. Benchmarks showed decode throughput of 41.8 tokens per second with PyTorch MPS BF16, rising to 134.8 with MLX BF16 and 188.9 with MLX INT4. The INT4 quantized model was 4.52 times faster than the PyTorch MPS baseline and produced a file 71.8% smaller than its BF16 equivalent. However, 4-bit quantization altered the generated text output in two out of three test prompts, indicating a trade-off between speed and output fidelity. The MLX delegate, which routes PyTorch computation graphs to Apple's MLX framework for GPU execution, remains experimental and subject to API changes.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in