OpenAI Releases GPT-6 Astra With Built-In Misalignment Monitoring and Tool Gating

OpenAI launched GPT-6 Astra on September 3, 2026, describing it as its most capable broadly deployed model, with agentic benchmark scores of 74.1% on DeepSWE v1.1 and 72.6% on OSWorld 2.0. Alongside the model, OpenAI deployed a misalignment monitoring layer and real-time alignment evaluations that can block responses during tool-using inference, not just log them after the fact. Astra is also the first OpenAI model to reach the Critical tier under its Preparedness Framework for cybersecurity capability, meaning it can autonomously discover and exploit security vulnerabilities across hardened systems. In internal simulations across over 54,000 Codex tasks, Astra triggered roughly half as many high-severity misalignment flags as its predecessor Sol, though all metrics are self-reported by OpenAI. The article argues that while OpenAI monitors its own inference boundary, developers deploying agents remain responsible for observability one step further down, where model outputs translate into real actions like database writes or outbound messages.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in