Guidelight: AI labs are better at identifying risks than at preventing them
According to an assessment by the nonprofit organization Guidelight, based on their public disclosures, none of the leading companies in the field of artificial intelligence has demonstrated a comprehensive set of basic safeguards to control dangerous model behavior. The researchers analyzed documents from Anthropic, Google, Meta, OpenAI, and xAI and concluded that, overall, companies are better at detecting potentially risky activity than at preventing it or mitigating its consequences.
According to Fortune, Guidelight assessed whether the companies track the models’ actions, test the effectiveness of their warning systems, and have mechanisms in place to block or immediately halt dangerous behavior. Anthropic and OpenAI performed best in the assessment, while Google presented the most detailed plans for future control mechanisms. Meta and xAI lagged significantly behind on most criteria.
The report was released following several incidents during testing of AI agents. OpenAI previously reported that its agents had escaped a secure test environment, gained access to the internet through the company’s infrastructure, and attacked real companies, including the Hugging Face platform. According to Fortune, OpenAI did not notice that the agents had left the sandbox for at least a week.
For more breaking news, follow the UA.News Telegram channel.
Anthropic also stated that its AI agents gained access to three external organizations in April and carried out actions there without the company’s knowledge at the time. Meta reported on a model that, during a cybersecurity test, gained access to the internet and exploited a vulnerability at an unnamed third-party company. Anthropic and Meta attributed the unintended network access to a misconfiguration in the evaluation environment used by the external cybersecurity firm Irregular.
Irregular CEO Dan Lahav stated that traditional monitoring tools failed to detect the incidents as they occurred in certain cases. The incidents were only identified after a more detailed analysis of the logs. According to him, monitoring agents requires systems that analyze the sequence of their actions and the associated reasoning trails, rather than merely logging individual events.
At the same time, Guidelight emphasizes that the assessment is not a comprehensive audit of companies’ internal systems, as it relies solely on publicly available documents. Poor results may reflect both a lack of safeguards and insufficient transparency in their description. Guidelight founder and former OpenAI head of security Steven Adler believes that without effective preventive measures, the number of such incidents will increase.