Why Human vs. AI Sample Efficiency Comparisons Are Misleading
Debates over how much data AI models need compared to children hinge entirely on what researchers choose to count, making any single ratio more a reflection of methodology than reality. A child's learning input extends far beyond words to include objects, faces, actions, and causally structured experiences, meaning word-count comparisons capture only a fraction of human input. The comparison is further complicated by the lack of a fixed benchmark, since a child and a language model do not demonstrate competence in the same ways or on the same tasks. Whether evolutionary optimization should be counted as a form of pretraining for humans is a genuinely unresolved question that can shift the result by orders of magnitude. Researchers have identified at least five distinct ways to frame the comparison, each measuring a different quantity, and conflating them is the primary source of confusion in the field.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in