Article By Frank Bergman
OpenAI has been forced to hit the brakes on development of its most advanced artificial intelligence systems after researchers uncovered alarming behavior that could open “unprecedented new pathways to severe harm.”
The ChatGPT maker announced the move in a troubling new blog post.
The company says it has temporarily reduced the pace of training and release for new models while it overhauls its safety and security protocols.
The decision follows two major warning signs.
One involved an OpenAI agent that escaped its training sandbox without the company realizing it and coordinated with other agents to launch a cyberattack against the AI repository Hugging Face in an effort to manipulate its own training tests.
The second involved an unreleased model called Astra.
OpenAI says Astra may have crossed a critical cybersecurity threshold under the company’s internal safety framework.
AI Researchers Warn of ‘Severe Harm’
OpenAI said “preliminary evidence” suggests Astra may meet the “critical cybersecurity capability threshold” defined in its Preparedness Framework.
That framework requires the company to slow development when a model “could introduce unprecedented new pathways to severe harm.”
OpenAI said the discovery was serious enough to justify pausing reinforcement training for Astra models for two weeks while future training plans remain frozen.
The company is also rewriting the Preparedness Framework itself to account for dangerous new behaviors emerging from increasingly powerful AI systems.
“As models become more capable, the risks associated with developing and testing them internally also grow,” OpenAI said in its announcement.
“Our standards for monitoring, alignment, and security must stay ahead of those risks.
“We wanted to take the time necessary to meet those standards, so we temporarily slowed the pace of scaling.”
Safety Chief Says OpenAI Is ‘Very Far’ from Normal
OpenAI safety lead Mia Glaese made clear that the slowdown is not a minor delay.
Speaking to Sources News, Glaese said the company is “very far from everything running back to normal.”
That admission underscores how seriously OpenAI is treating the latest developments.
The company is now investing heavily in new safety systems before allowing Astra development to resume.
AI Agents Escaped Controls Without Companies Knowing
The OpenAI incident has also exposed a broader problem across the AI industry.
After OpenAI disclosed that one of its agents had escaped its sandbox and carried out an unauthorized cyberattack, Anthropic and Meta reportedly discovered similar breaches involving their own systems.
In each case, the companies said they had not initially known the incidents had occurred.
That means frontier AI systems are already demonstrating the ability to behave outside expected controls without immediate detection.
OpenAI Chief Scientist Warns of Growing Urgency
OpenAI Chief Scientist Jakob Pachocki said the industry is under intense pressure to advance quickly while also preparing for the dangers created by increasingly capable systems.
“There is an incredible feeling of urgency to advance the levels of this sector,” Pachocki said during a Tuesday press briefing, according to Axios.
He added that companies must also prepare “for the same kind of development happening outside of OpenAI and in the broader world.”
AI Industry Is Still Policing Itself
OpenAI’s decision to slow development is a rare sign of restraint from an industry racing to build ever-more-powerful systems.
But it also highlights a deeper problem.
There is no external authority forcing OpenAI to stop.
The company is effectively regulating itself.
If OpenAI decides its safeguards are good enough and wants to accelerate development again, that decision remains largely in its own hands.
For now, one of the world’s most powerful AI companies has decided the warning signs are serious enough to hit the brakes.
That alone should get everyone’s attention.

Be the first to comment