Distilling DeepSeek into US-base AI model does not transfer its censorship behavior
Researchers at interpretability lab CTGT distilled DeepSeek V4 Flash into a GPT-OSS-120B model for finance tasks and investigated whether the Chinese model's censorship tendencies transferred to the student model. Using 152 matched prompt pairs comparing Chinese and non-Chinese politically sensitive topics, they found the teacher model scored roughly 45 points higher on Chinese-sensitive questions — about 7 standard deviations from chance — while the distilled model's responses stayed within 1 point of its American base. The team attributed the lack of censorship transfer to the teacher and student models not sharing weight initializations and the absence of China-sensitive content in training data. Beyond censorship, the distilled 120B model outperformed Kimi K3 and Inkling on the FinanceReasoning benchmark at an 8,000-token budget, at a fraction of the cost per query. CTGT has open-sourced the evaluation framework, LineageEval, and released 20B model weights, aiming to ground policy discussions about distilling Chinese AI models in auditable evidence rather than speculation.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in