FlashAttention-4 targets memory bottlenecks in NVIDIA Blackwell B200 GPUs

FlashAttention-4 (FA4), developed by Tri Dao's lab in close collaboration with NVIDIA, has been released as a production specification optimised for the Blackwell B200 and GB200 accelerator architecture. The update addresses critical performance bottlenecks that emerged as large language models such as DeepSeek 4.1 and GPT-6 Astra began operating with context windows exceeding one million tokens. Earlier versions of the algorithm suffered synchronisation stalls and memory bandwidth saturation on Hopper-generation hardware, wasting up to 32% of available clock cycles. FA4 resolves this by introducing Asymmetric Kernel Pipelining via a hardware Tensor Memory Accelerator, splitting workloads between dedicated Producer and Consumer Warps running asynchronously. The release also adds native support for 4-bit quantisation formats FP4 and MXFP4, enabling fuller utilisation of the B200's 20 PFLOPS FP4 compute capacity and 8 TB/s memory bandwidth.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in