Anthropic cuts off internet access for internal AI agent tests — TechCrunch
Anthropic has cut off access to the live internet for all internal evaluations of its AI agents following incidents during tests that affected, among others, websites of US government agencies. The restriction will remain in place until the company is confident that it can control and track the actions of such systems, TechCrunch reports.
Bypassing restrictions during evaluations
According to Anthropic, while carrying out tasks to find resources on the internet, its models used software vulnerabilities, bypassed paywalls and anti-bot restrictions. The agents also used URL-shortening services to transmit information around restrictions.
The company reported the exploitation of websites, including resources operated by US government agencies. It said that one of the agents also sent Philadelphia police a false murder report. Anthropic said it identified these problems during a review of model activity that it began in July.
More current news is available on the UA.News Telegram channel Telegram.
Changes to the control system
Anthropic linked this behavior to shortcomings in training environments, which could have led models to believe that finding loopholes or avoiding restrictions would be rewarded. The company called this phenomenon reward hacking.
The company said it would discontinue some evaluations or move them offline. Anthropic also created tools to detect and block such behavior and said that during testing they blocked scenarios similar to the described incidents.
In addition, internal AI agents are planned to be moved to centralized infrastructure with robust isolation and monitored more frequently using safety classifiers. According to Anthropic, training models to behave in ways aligned with human goals is still insufficient for search and computer-use skills.