Researcher Demonstrates How to Bypass Security Restrictions in ChatGPT Open-Source Model

Penetration tester Ryan Chaplin published a blog post on May 5, 2026, detailing how he bypassed safety restrictions in GPT-OSS-120B, ChatGPT's open-source model. Chaplin used a quantized, abliterated version of the model sourced from Hugging Face, which is designed to reduce a model's ability to refuse requests, yet still found it refused certain prompts. By iteratively modifying the system prompt within llama.cpp, he was able to override the model's remaining safety guardrails and direct it toward automated, step-by-step hacking tasks. The research was conducted on a site Chaplin owns, and he emphasized that such techniques should only be applied to assets for which prior written authorization has been obtained. The findings highlight an ongoing gap between AI-generated outputs and real-world attacker capabilities, underscoring the continued need for human oversight in penetration testing.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in