Speculative Decoding Speedups Depend More on Acceptance Rate Than Hardware
A developer investigating slower-than-expected speculative decoding performance in LLM inference found that Apple Silicon's MPS backend was initially blamed for the gap due to high fixed dispatch costs per model call. However, hardware-agnostic FLOPs-based theoretical calculations using Leviathan et al.'s formula revealed the algorithm itself could not surpass 1.0x speedup at the observed acceptance rates, even with zero hardware overhead. Switching from a 0.5B to a 1.5B draft model — expected to improve distribution matching with the 3B verifier — produced highly variable results, with acceptance rates swinging from 19% to 86% across runs on the same prompt type. The investigation concluded that acceptance rate, driven by prompt characteristics and draft-verifier distribution alignment, is the primary ceiling on speculative decoding gains. Hardware overhead compounds the problem but is not the root cause of underperformance.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in