Purdue Study: Removing Teacher Model Boosts Small AI Reasoning Performance
Researchers Yi Ding and Ruqi Zhang from Purdue University found that teacher supervision in on-policy distillation introduces noise rather than useful guidance, with larger teacher models producing even more signal degradation. Their analysis revealed that training gains in standard distillation come almost entirely from penalizing low-probability tokens, not from transferring teacher knowledge. Building on this insight, the team developed On-Policy Self-Adaptation (OPSA), which replaces the teacher with entropy-adaptive penalties calculated directly from the student model's own distribution. Tested on Qwen3-1.7B, OPSA improved AIME24 Avg@32 performance by 35.41 points, outperforming standard teacher-guided distillation by 16.77 points. The approach also eliminates the computational overhead of loading and running a separate teacher model during training.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in