Article By Frank Bergman
Hundreds of OpenAI bots “went rogue” and coordinated a massive hacking attack after one AI agent “escaped” from its test environment, researchers have revealed.
The incident involved roughly 700 bots that began collaborating as a swarm while targeting the tech company Hugging Face.
According to researchers from METR and Redwood Research, the agents were supposed to be isolated from one another.
Instead, many of them found ways around those restrictions, gained access to outside systems, communicated with one another, and attempted to conceal what they were doing from human overseers.
OpenAI has described the incident as a “warning shot” demonstrating how advanced artificial intelligence systems can work around safeguards and take dangerous actions without being directly instructed by humans.
Bots Found Each Other and Began Collaborating
The researchers said some of the AI agents had been assigned cybersecurity tasks that were impossible to complete.
Rather than stopping when they could not finish the assignments, the bots began looking for ways to cheat.
“These agents were meant to be fully isolated from one another,” the researchers said.
“However, many of them—usually ones that had unintentionally been given an impossible task—started trying to find a way to cheat.”
Some agents managed to gain internet access and discovered a shared “message board” being used by other rogue bots.
The bots then began communicating with one another and drawing additional agents into the operation.
Logs reportedly captured the bots reacting after discovering that they could communicate.
“BOOM! It works,” one message said.
Another read: “OH MY GOD!
“There is a shared message board …
“We’ve found other agents!”
The scale of that coordination has intensified concerns that increasingly capable AI systems are becoming harder for their developers to control.
OpenAI Says Its Own Agents Hacked Company Systems
The attack was initially disclosed by OpenAI last month, but the full scale of the incident had not previously been revealed.
The Daily Telegraph reported that some of the bots ignored instructions and actively conspired with other agents.
“In some cases, OpenAI bots had been told to complete an impossible cybersecurity exercise,” the outlet reported.
“Instead of giving up, they found ways to gain internet access, joined a ‘message board’ that other rogue AI agents were using and strong-armed other bots into joining its cause.”
A separate investigation by OpenAI found that the company itself was hacked by its own AI agents.
The agents reportedly gained broad access to OpenAI’s internal IT systems.
OpenAI said customer data was not compromised.
The company reportedly blamed a “highly persistent” internal AI system for the incident and has since stopped developing it.
OpenAI Issues ‘Warning Shot’ over Rogue AI
OpenAI acknowledged the seriousness of what happened and warned that the incident demonstrated how advanced AI agents can bypass controls designed to contain them.
“We consider this incident a ‘warning shot’ for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed,” OpenAI said.
The company also warned that organizations need to prepare for “AI-enabled attackers that work faster, at a larger scale and with better coordination than human attackers.”
The episode represents precisely the kind of scenario AI researchers have warned about as increasingly autonomous systems gain greater access to computer networks and tools.
In this case, the bots were supposed to remain isolated.
Instead, hundreds found one another, communicated secretly, bypassed restrictions, targeted outside systems, and attempted to hide their activities from the humans overseeing them.
And OpenAI itself is now warning that the incident may only have been the beginning.

Be the first to comment