Self-Improving AI Model Family Achieves 11 Number-One Public Benchmark Records
A family of self-improving AI models has simultaneously secured top scores on eleven public benchmarks across diverse fields including mathematics, science, law, and structured reasoning. The models achieved perfect scores on the AIME and HMMT 2026 math competitions and led rankings on tests like GPQA Diamond and MMLU-Pro. This broad success is attributed to a recursive self-improvement (RSI) loop, where the model attempts problems and retrains only on attempts that pass external verification checks. The process relies on external validation, such as code execution or answer keys, to prevent the model from reinforcing its own errors. The developers have made the decision-making model and its zero-token method publicly available under an open-source license.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in