OpenAI says it is adding new security safeguards and pausing development after internal testing found that an upcoming model crossed a "critical" capability threshold, according to CPO Magazine.
The model in question is named Astra. Per CPO Magazine's reporting, OpenAI's internal findings indicate Astra has crossed "critical cybersecurity capability" thresholds — the company's own internal marker for a model whose abilities are dangerous enough to require intervention before release.
In response, OpenAI is promising enhanced security safeguards and a pause on development for a period the company describes as "reinforcement." In other words, work on the model stops while protections are strengthened around it.
What makes this notable is that the trigger came from OpenAI's own evaluations rather than from an outside researcher, regulator, or incident. AI developers have spent the past few years publishing internal frameworks that define capability tiers and commit the company to specific actions when a model reaches a given tier. Those commitments are largely voluntary and self-policed, which is why an announced pause is closely watched: it is a test of whether such a framework actually binds a company's behavior when a model gets too capable.
"Critical cybersecurity capability" generally refers to a model's ability to meaningfully assist in offensive computer operations — the sort of skill that would be valuable to attackers as well as defenders. The source items do not detail what Astra can specifically do, how long the pause will last, or what the added safeguards involve.
It matters because a leading AI lab has publicly stopped its own product over a security risk it found itself, offering a rare, if incomplete, look at where the industry is drawing its safety lines.