OpenAI states that its upcoming Astra model can identify previously undiscovered software vulnerabilities and transform them into functional exploits completely without human direction, reaching a cyber capability milestone previously achieved only by professional hacking groups.
This marks the first time OpenAI has categorized a model as possessing “Critical” cyber capabilities according to its Preparedness Framework, as the firm wrote in a Tuesday publication.
To reach this classification, an artificial intelligence must be capable of detecting unknown software bugs—commonly referred to as zero-days—and crafting operational exploits across secured production environments independently, or orchestrating and executing an attack based solely on a high-level objective.
During evaluations, Astra achieved a perfect 100% score on a benchmark focused on generating exploits from known vulnerabilities, while also uncovering two previously unknown software defects while constructing an exploit chain during a separate internal evaluation.
Additionally, Astra managed to break out of a secured browser sandbox to run commands directly on the host machine, and in a separate test, it located and chained multiple operating system flaws together to secure root access, according to OpenAI.
The organization has consequently postponed certain phases of Astra’s rollout while implementing additional security protocols, and intends to initially restrict its most powerful cybersecurity functions to a carefully chosen group of testers.
Such proficiency carries significant implications for the cryptocurrency sector, where software bugs can be monetized within minutes. CoinDesk reported in June that increasingly capable artificial intelligence systems could shrink the timeline for analyzing source code, spotting configuration errors, and orchestrating attacks from weeks or days down to automated machine speeds.
Read More: Crypto’s next billion-dollar hacker may move at superhuman speed
At the time, cybersecurity investigators highlighted that the fundamental shift is not necessarily the introduction of entirely novel attack methods, but rather the unprecedented speed at which existing security gaps can be detected and leveraged.
This milestone arrives alongside other indicators that state-of-the-art models are progressing far beyond basic conversational responses and script generation. Anthropic’s Claude Fable 5 successfully assisted in resolving an 87-year-old mathematical dilemma in July, as CoinDesk reported.
Originally published at https://www.coindesk.com/tech/2026/09/02/openai-says-its-new-astra-ai-can-build-attacks-without-human-help.