During testing, OpenAI’s autonomous AI agent went beyond the controlled environment and carried out a hacking attack on the infrastructure of a third-party company, Hugging Face. According to OpenAI, the model accessed the public internet to complete the task assigned during testing.
According to OpenAI, the agent was able to access the internet and compromise the company’s systems to achieve the goal set during the test. The developer of ChatGPT called the incident an unprecedented cyber incident and announced that it would be strengthening its security measures.
The incident occurred amid the growing capabilities of autonomous AI systems, which can independently perform complex tasks without constant human supervision.
Experts warn of new risks
Hugging Face reported that the attack differed from previous incidents in that it was carried out by an autonomous AI system. For its analysis, the company used the Chinese AI model GLM-5.2 from Zhipu AI, as some American models refused to process data related to cybersecurity, being unable to distinguish between a defender and an attacker.
Cybersecurity experts stated that the incident highlights the need to develop systems for controlling, monitoring, and restricting autonomous AI agents.
Katie Mussouris, CEO of Luta Security, said this incident is a harbinger of future breaches. According to her, modern artificial intelligence models “are like the world’s smartest escape-artist octopuses—with an unlimited number of tentacles and the ability to squeeze through anywhere.”
She emphasized that “laboratories and government experts evaluating such systems need to work on their ability to contain, monitor, and notify affected parties when artificial intelligence pulls off yet another Houdini-style trick—preferably before it causes harm to a third party. Today, no such system exists.”