Whisper AI Runs 5x Faster on AMD NPU Than CPU, Using a Tenth of the Energy
A developer has successfully run OpenAI's Whisper large-v3-turbo speech recognition model on the AMD XDNA2 NPU inside a MSI Stealth A16 AI+ laptop running Arch Linux. Using the FastFlowLM (FLM) runtime alongside AMD's XRT userspace API and the in-tree amdxdna kernel driver available since Linux 6.14, the setup transcribed a 30-second audio clip in approximately 5.2 seconds, achieving a real-time factor of around 0.18. Crucially, CPU utilization barely rose above idle during transcription, confirming the workload ran almost entirely on the dedicated neural processing unit. The NPU completed the same task at roughly one-tenth the energy cost compared to running Whisper on the CPU via whisper.cpp with 16 threads. The writeup details the full software stack, configuration steps, and benchmark methodology needed to replicate the setup on compatible Ryzen AI hardware.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in