OpenAI's GPT-6 Astra Rated 'Critical' for Cyber Risk as Prompt Injection Threat Persists

OpenAI launched GPT-6 Astra on September 3, 2026, marking it as the first of its models to reach the 'Critical' cybersecurity capability level under the company's own Preparedness Framework, meaning it can autonomously discover and exploit unknown security flaws across hardened systems. The model's system card, citing external evaluations, reports a prompt-injection attack success rate of 8.5%, down from 27% for its predecessor Sol, though security researchers note that improvement is not the same as elimination of the threat. At scale, even an 8.5% success rate across thousands of daily tool-using agent sessions — where attacker-controlled text can enter via uploaded files, emails, or third-party APIs — translates into a meaningful real-world risk. Third-party evaluator Artificial Analysis ranks Astra 14th among 202 tracked models with an Intelligence Index of 60, offering a more measured view than OpenAI's own benchmark figures. Developers building agentic systems are being urged to treat prompt injection as a structural architecture problem rather than a model-level issue, since no lab has yet published a complete fix for the vulnerability class.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in