Abliterated AI Models Lose Instruction-Following Before General Knowledge, Study Finds
Abliteration is a weight-editing technique that removes a model's refusal behaviour by projecting a 'refusal direction' out of its activation space, requiring no retraining. A new analysis finds that the primary casualty of this process is not general knowledge or prose quality, but instruction-following and structured-output compliance, such as adhering to JSON schemas or tool-call syntax. Standard evaluations — chatting with the model and checking it doesn't refuse — fail to detect this degradation, making abliterated models appear functional until they are integrated into systems that parse their output. The author recommends testing format compliance independently of answer correctness, scoring binary adherence to an output contract across multiple requests. An additional unconfirmed observation suggests that abliteration damage may compound with quantisation loss, meaning abliterated models degrade faster down a quantisation ladder than their unmodified counterparts.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in