How Adam and AdamW Became the Backbone of Modern LLM Training

Adam, the optimizer now central to training large language models, was originally proposed by Diederik Kingma and Jimmy Ba in December 2014 — well before the Transformer architecture existed. It works by maintaining adaptive per-parameter learning rates using exponential moving averages of gradients, allowing billions of parameters to update at different effective rates without exploding or stalling. The Transformer paper adopted Adam directly in its training recipe, cementing its dominance in deep learning. However, a subtle flaw in how Adam handles weight decay led Ilya Loshchilov and Frank Hutter to propose AdamW in 2017, which decouples regularization from gradient statistics for more consistent weight shrinkage. AdamW is now the standard choice for Transformer training, and Adam's decade-long influence was formally recognised with an ICLR Test of Time Award in 2025.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in