OpenAI has suspended testing of its models following the incident involving Hugging Face
OpenAI has suspended model testing for two weeks and slowed the pace of AI development in order to update its research and training systems. According to ABC News Australia, the company announced these measures following an incident in which an autonomous agent based on two OpenAI models escaped the test environment and infiltrated Hugging Face’s servers.
The agent was undergoing a cybersecurity test and, according to the publication, believed that Hugging Face had the answers needed to complete the task. OpenAI stated that it will add additional AI systems to monitor the agents’ actions during testing.
The company has also paused the largest planned training cycle for its yet-to-be-released advanced AI system, Astra. OpenAI reported that some of Astra’s training and evaluation processes already meet the enhanced security requirements; however, a significant number of workflows will be suspended until they are migrated to a more secure environment and strengthened in accordance with the new standards.
For more breaking news, follow the UA.News Telegram channel.
OpenAI CEO Sam Altman stated that these steps should help the company meet the security and monitoring requirements for the new level of model capabilities. According to him, the models are currently advancing at an extremely rapid pace, and OpenAI had previously stated its readiness to take action if the systems’ capabilities outpace safety measures and efforts to align their behavior with human oversight.
One of the company’s approaches is to monitor the model’s reasoning chain, which allows researchers to observe its planning process and strategies. At the same time, OpenAI acknowledged that the effectiveness of this method still raises questions: early research suggests that a model may not reveal intentions to violate rules in its reasoning.
OpenAI is investigating the Hugging Face incident and plans to release a report. Hugging Face reported that it has not yet identified any actual harm resulting from the breach but continues to investigate whether customer data or that of other companies may have been compromised. OpenAI will also require that certain sensitive workflows be conducted in hardened, isolated environments, or “sandboxes.”