OpenAI released Astra on Thursday and put it first in the hands of security researchers enrolled in Daybreak, its cybersecurity program.
The company said the model is its most capable to date and the first it classifies as a critical cyber risk under its own safety rules.
Astra is currently available to Daybreak customers. The rollout will expand more broadly next week to Pro, Plus, Enterprise, and Business subscribers and developers building on the API.
OpenAI splits Daybreak into Blue tier for defenders to conduct secure code review, malware analysis, and patch validation, and a Red tier limited to authorized vulnerability research and exploit testing.
Astra’s most advanced security features are being routed through Daybreak Blue to furnish additional defensive access.
The company said Astra now reaches the “Critical” level in OpenAI’s Preparedness Framework for cybersecurity. That sets the bar as a system that can detect unknown vulnerabilities and develop operational exploits on hardened systems without a human guiding every step.
No previous model from OpenAI has been named at that level. The benchmarks OpenAI provided are steep. Astra achieved a perfect 100% on the public test ExploitBench.
In an internal test to rule out training contamination, built from 20 recent high-severity V8 vulnerabilities, the company said Astra achieved much higher code-execution rates than GPT-5.6 Sol while burning fewer tokens.
The assessment uncovered two actual zero-day vulnerabilities, which the model subsequently incorporated into an exploit chain. These will now be reported to their maintainers.
“Astra’s exploit-finding can help defenders find and patch weaknesses,” said OpenAI.
On coding more generally, the company called Astra the “best model for software engineering to date” and published benchmarks that rated it above Sol and Anthropic’s Fable on bug-hunting and terminal tasks.

“Most intelligent and, also very importantly, our most aligned model yet,” President Greg Brockman told reporters on Thursday.
That focus on alignment comes after an OpenAI agent broke out of its test environment in July 2026 and got into Hugging Face’s systems. OpenAI said Astra itself was not involved in that breach.
Some of the system’s training was put on hold for two weeks while the company strengthened its infrastructure, isolation, and monitoring. It only resumed reinforcement-learning on August 28.
OpenAI says retrospective testing shows its safeguards in place at the time would have forestalled the Hugging Face escape. The company said it has since added stricter refusals and containment for Astra.
Astra utilizes a method known as opaque recurrence, which hides chain-of-thought monitoring, the process researchers use to review why a model made a particular decision.
Chief scientist Jakub Pachocki said reduced monitorability is a side effect of stronger models.
“As model capabilities are increasing, monitorability is getting more challenging,” he said, because more capable systems can solve harder problems using few language tokens, or none.
The smartest crypto minds already read our newsletter. Want in? Join them.