Study Finds No Single Mechanism Causes Neural Network Training Loss Spikes
A controlled experiment testing four competing theories about loss spikes in neural network training found that none of the proposed mechanisms — edge-of-stability thresholds, weight-norm collapse, numerical precision, or non-normal amplification — is sufficient on its own to trigger spikes. The study used a standardized MLP benchmark across 60 runs to measure each mechanism's proposed trigger simultaneously, something prior 2026 papers had not done. Results showed that weight-norm collapse occurred in runs with zero spikes, and crossing the sharpness threshold was statistically no better than a coin flip at predicting spikes. A freezing experiment revealed that spikes require both a sharp training regime and ongoing adaptation of scale-invariant hidden projection weights — remove either condition and spikes cease. The findings, currently under peer review, are limited to a toy MLP under plain SGD and have not yet been validated in larger transformer or AdamW settings.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in