Vocabulary Caps Hurt Small Language Models, New Experiment Finds
A developer running experiments on small transformer models tested whether capping vocabulary size — a feature associated with the TinyStories paper — improves model performance on a cs.CL abstracts corpus. Three vocabulary caps (1,500, 4,000, and 8,000 types) were tested on an identical 4-layer model with 846K training tokens, scored only on token positions common to all arms to ensure fair comparison. The 1,500-token cap scored 4 accuracy points lower than larger vocabularies and produced severely degraded text, with over a third of output tokens replaced by an unknown-word placeholder. Unlike TinyStories, where simple words carried full meaning in children's stories, capping vocabulary on a technical corpus strips content words and forces the model into repetitive, skeleton-like outputs. The experiment concludes that TinyStories' success came from its domain and writing style, not from the vocabulary restriction itself.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in