Study finds lack of public AI containment plans
Leading developers of artificial intelligence systems have few publicly documented action plans for cases in which a model attempts to bypass human control. This was reported by TechCrunch, citing an assessment by Guidelight AI Standards.
Guidelight analyzed public materials from Anthropic, Google, OpenAI, Meta, and xAI against six priority practices in its Control standard. Among other things, the organization assessed internal logging and monitoring of AI actions, halting systems after a series of signals of improper behavior, independent audits of controls, and the existence of a specific model containment plan.
Assessment of public readiness
A containment plan provides for predetermined actions after an attempt by AI to bypass control is detected. It should specify which permissions need to be revoked from the model, for whom it may continue operating and under what restrictions, as well as when it should be fully shut down.
OpenAI received the highest score — three out of five. According to Guidelight, the company has repeatedly paused or terminated work processes, including internal deployment and model training, following security incidents. The company also described the steps required to resume these processes.
At the same time, the Guidelight report states that there is no public evidence that OpenAI has a formalized response plan for future incidents involving models becoming misaligned with human control.
More current news is available on the UA.News Telegram channel Telegram.
Meta and Anthropic received the lowest scores specifically for disclosing containment plans. Guidelight found no confirmation that Meta has a response plan for losing control of a model or intends to introduce one. In Anthropic's August risk report, according to the organization's assessment, there is no mention of limiting model deployment as a possible outcome of investigating incidents involving misalignment or control.
Company positions and regulatory requirements
An OpenAI representative said the assessment does not cover all of the company's internal practices. According to him, OpenAI has processes for restricting permissions, pausing workloads, narrowing deployment, or fully shutting down a model, and has already used them. Google also said that the report does not reflect the full range of its safety measures. Meta declined to say whether it has an internal containment plan, citing its existing risk assessment framework.
Guidelight's chief scientist and former OpenAI safety researcher Steven Adler said he was surprised by how little companies disclose about responding to a serious loss-of-control incident. In his view, without a prepared plan in advance, developers will have to determine actions during an emergency.
California's SB 53 law, which took effect this year, requires major developers of frontier models to publish frameworks for detecting and responding to critical safety incidents. In New York, the RAISE Act with similar requirements is set to take effect in January. Last month, a bipartisan AI Kill Switch Act bill was also introduced in the United States, requiring major developers to create and maintain technical mechanisms for shutting down uncontrolled models.