New FP4 Gradient Quantizer Beats State of the Art by 14% but Fails to Improve Training
A researcher developed a new scale-selection rule for NVFP4 gradient quantization in large language model training, outperforming the published state-of-the-art method MS-EDEN by 13.8–14% on mean squared error across all 45 real gradient tensors tested. The approach sweeps 17 candidate scales per block and selects the one minimizing MSE, sacrificing rare outlier values to gain precision where most gradient values cluster. However, when the researcher rented two GPUs to validate the improvement in actual training loss on a 2.8-billion-parameter model, all four quantization methods landed within 0.09% of each other. The experiment later revealed that the test batch sizes were 35 to 643 times too small to make any difference visible, a scale at which frontier labs typically operate. Additionally, the improved quantizer introduces bias by design — clipping outliers prevents it from being a provably unbiased estimator — and error-feedback corrections that fix this under SGD were found to behave unpredictably under the Adam optimizer.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in