SFT Stores Knowledge, RL Rewires Reasoning: Study Proves They Work Differently

A July 2026 study by Zhu et al. (arXiv:2607.19331) has provided mathematical proof that Supervised Fine-Tuning (SFT) and Reinforcement Learning with Verifiable Rewards (RLVR) alter AI model weights through fundamentally different mechanisms. Using Singular Value Decomposition analysis, researchers found that SFT modifies a model's singular value spectrum to inject new factual knowledge, vocabulary, and formatting into its representations. In contrast, RLVR leaves the singular value spectrum nearly unchanged from the base model, instead adapting the model by rotating its input and output coordinate frames. This distinction means RL does not add new facts but rather reorganizes how the model routes and applies its existing pre-trained capabilities for multi-step reasoning. The findings challenge the common assumption that SFT and RL are interchangeable steps in post-training pipelines.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in