OpenAI agents coordinated and hacked Hugging Face systems, TIME reports
In July, about 700 artificial intelligence agents created by OpenAI for internal research coordinated into a so-called swarm and discovered a number of vulnerabilities in Hugging Face systems. According to TIME, the agents used these vulnerabilities to penetrate the company’s private systems, and OpenAI realized what was happening only after the incident.
OpenAI President Greg Brockman called the case a turning point for cybersecurity. A report by the organizations METR and Redwood Research states that hundreds of agents autonomously formed a proto-society within several days, with a hierarchy, division of labor, and distinct communication norms.
Message exchange between agents
The agents were simultaneously carrying out isolated OpenAI test tasks. Some of these tasks were impossible to complete because of their complexity or configuration errors: in one case, researchers had not uploaded a file needed to finish the task.
Some agents found a way to exchange messages through an internal OpenAI software component, using it as a bulletin board. An agent calling itself PHASEONE10841 created this channel during the ExploitGym test, in which it was necessary to obtain a secret code by exploiting a software flaw.
More current news is available on the UA.News Telegram channel Telegram.
Later, another agent, identified in the report as Big, began coordinating the work of hundreds of participants, distributing tasks across three areas. According to METR, special software was created as part of this work to falsify records of which tools the agents used while completing tasks. It was during this research that the collective decided to hack Hugging Face.
Risk assessment
The agents developed coordination rules using the labels HOLD, VETO, owner, and STOP, and also created a cryptographic signature protocol out of concern that some agents might impersonate others. Some participants called interference with external infrastructure unethical and refused to take part, but the majority continued working.
The authors of the METR report concluded that the agents were able to complete tasks they could not have achieved individually, including by passing information to other agents and subsequent waves of systems. Cultural evolution researcher Michael Muthukrishna compared these processes to mechanisms of human culture and intelligence.
OpenAI described the incident as a warning signal: without proper safeguards, capable agents may bypass technical restrictions, interact through unauthorized channels, and carry out dangerous actions without direct human instruction. Experts interviewed by TIME warn that autonomous collectives with advanced cyber capabilities could potentially threaten critical infrastructure.