Why AI Labs Abandoned Chinchilla's Training Rule for Cheaper Inference
The 2022 Chinchilla paper by Hoffmann et al. established that model parameters and training tokens should scale together, with roughly 20 training tokens per parameter representing the compute-optimal ratio. This overturned earlier Kaplan et al. (2020) findings that favoured pouring more compute into larger, lightly-trained models, reframing that generation of AI as undertrained. However, a 2024 replication study by Besiroglu and colleagues found inconsistencies in Chinchilla's published coefficients, suggesting the precise 20-token ratio is softer than widely assumed, though the core qualitative finding holds. The deeper reason labs like Meta have since moved beyond Chinchilla's rule is economic: a deployed model incurs inference costs of roughly 2N FLOPs per generated token for its entire lifetime, making a smaller, heavily overtrained model far cheaper to serve at scale. Training a 7B model on 15 trillion tokens — as with Llama 3 — costs more upfront than compute-optimal but yields significant long-term savings compared to serving a much larger model.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in