Full Hierarchical Pitman-Yor Model Fails to Match Transformer Data Scaling
A research series testing whether language-model-like behaviour can emerge without neural networks reached a key milestone with its final non-neural candidate. The experiment pitted a fully trained hierarchical Pitman-Yor language model — using Gibbs sampling and inferred hyperparameters — against a small transformer on a text prediction benchmark. While the full model outperformed simpler count-based methods like modified Kneser-Ney, its accuracy gains halved with every doubling of training data, mirroring weaker models rather than closing the gap with the transformer. The transformer continued improving at roughly three times the rate of the best count-based model once training data exceeded one million tokens. The finding suggests the ceiling lies with the count-model family itself, not with how well its parameters are fitted, meaning any viable non-neural alternative will require a fundamentally different type of state representation.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in