OpenAI has paused part of the development of Astra, its upcoming AI model, after internal testing suggested the system may be approaching the company's own "Critical" cybersecurity threshold — a first for any OpenAI model.

According to The Tribune (India), OpenAI said on Friday that Astra had shown significant advances in agentic coding and cybersecurity, prompting the company to conclude it could not rule out the higher risk level. Forbes reported the pause under the framing that the model "nears" that first-ever critical rating.

What does "Critical" actually mean here? According to Deccan Chronicle, OpenAI's safety guidelines define the threshold as a model that can autonomously find and exploit severe, real-world software flaws — the kind known as zero-day exploits, because defenders have had zero days to patch them. A system that could do that on its own would compress work that today requires skilled human researchers, which is why the classification carries weight beyond OpenAI's internal paperwork.

The company is not shelving Astra entirely. Benzinga reported that CEO Sam Altman said Astra will be "generally available," but that its cyber capabilities require more safety work before release. In other words, the pause appears to apply to a specific slice of the model's abilities rather than the whole product.

Skepticism has followed. Techi.com noted that while Astra may be "Critical," the underlying proof remains private — OpenAI has not published the evidence behind the classification. Separately, 동아사이언스 and Yellow.com reported that OpenAI used Astra to tackle ten math problems, a demonstration that drew both praise and criticism. Times Now framed the episode as a model "the world may not be ready for yet."

Why it matters: this is the first time a leading AI lab has publicly held back part of a flagship model because it might be good enough at hacking to be dangerous — and, so far, we're being asked to take that on trust.