Student Builds 64M-Parameter Language Model from Scratch, Tests Reasoning Limits

A computer science student at Politeknik Negeri Jakarta built Klyra-64M, a 64-million-parameter language model trained entirely from random weights rather than fine-tuning an existing model. The project was motivated by curiosity about whether a very small language model — well under 100 million parameters — could still learn language, follow instructions, and perform basic reasoning. Using the open-source MiniMind-3 Dense architecture as a base, the model was pretrained on approximately one billion tokens from the Ultra-FineWeb dataset, achieving a validation perplexity of 10.643. The student then continued training with an additional two billion tokens to broaden the model's knowledge before moving on to instruction tuning and reinforcement learning experiments. The base checkpoint, Klyra-64M-Base, has been publicly released on Hugging Face for further research.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in