Anthropic and OpenAI propose evaluators without the power to halt AI model development
Anthropic and OpenAI have supported involving external safety evaluators in work with frontier artificial intelligence models, but such specialists will not have independent authority to halt model training or deployment. This was reported by CNBC Top News.
Anthropic CEO Dario Amodei proposed giving third-party evaluators ongoing access comparable to that of internal risk-management teams. Under his plan, they would also be able to publish their findings without company editorial control, although with limited information redactions.
Access without veto power
Amodei compared this approach to banking supervision practices, in which regulators work inside large banks. Julie Andersen Hill, dean of the University of Wyoming College of Law and an expert in banking regulation, considers this comparison inaccurate. According to her, banking supervisors can require certain practices to be stopped, limit an institution’s growth, seek changes in management and, in extreme cases, close a bank.
The evaluators proposed by Anthropic will be able to study systems and report risks, but will not have legal tools to prohibit the development or release of a model. OpenAI has also committed to introducing a similar mechanism, but has not yet disclosed details of how permanently engaged evaluators would operate.
More current news is available on the UA.News Telegram channel Telegram.
Questions of independence
Albert Ziegler, head of AI at cybersecurity company XBOW, whose team received early access to models from Anthropic, OpenAI and other developers, confirmed that evaluators do not have veto power. He noted that testing can reveal whether a model is capable of completing a dangerous task, but assessing the risk of the entire system requires access to instructions, tools, permissions, safeguards and logs of attempts to carry out actions.
Deborah Raji, a researcher at the University of California, Berkeley, stressed that access to a system alone does not make an evaluator independent. In her view, there should be a body outside the company that determines auditors’ qualifications, the scope of their review and where findings should be submitted. Anthropic named the nonprofit Model Evaluation and Threat Research, or METR, with which it has already worked, as a potential evaluator.
METR says it does not accept monetary payments or donations from AI companies or their executives. At the same time, the organization acknowledged in its own report that some of its employees have close social ties with employees of AI companies and that it shares a research center with staff from some laboratories.
Experts also pointed to the absence of a rules-based system for frontier AI similar to banking regulation: such a system would define prohibited actions, the limits of evaluators’ discretion and the consequences of serious findings. In Hill’s view, access and the right to publish results do not provide the trust inherent in government oversight if the developer retains control over final decisions.