Anthropic estimates risk of humanity’s extinction from AI at over 10% — The Verge
Evan Hubinger, the head of one of Anthropic’s safety teams, said that he personally estimates the probability that artificial intelligence could lead to the deaths of all humans over the next decade at more than 10%. The Verge reports.
Researcher’s dismissal
Hubinger’s statement came after researcher Jacob Coxon wrote on X about being dismissed from Anthropic. According to Coxon, he left the company because he considered its approach to safety insufficiently strict.
Coxon, who previously trained AI systems for OpenAI, accused Anthropic and OpenAI of racing to create self-improving superintelligence. He said that the companies are “playing with our lives” and that people creating AI sincerely consider the possibility of humanity’s extinction by the end of the decade.
More current news is available on the UA.News Telegram channel Telegram.
In his response, Hubinger agreed with this characterization of the concerns. He also noted that the development of self-improving AI is progressing faster than expected.
Risks of self-improvement
The industry has long discussed the potential threats posed by AI systems capable of improving themselves independently. The concerns are linked to the possibility that such a process could get out of human control in a cycle known as recursive self-improvement.
As The Verge notes, such systems have not yet been created, but companies are actively seeking to develop them, and a significant share of modern software code is already written with the help of AI. According to Hubinger, Anthropic currently has no plan that would ensure the safety of advanced AI systems and their alignment with human values, and the company has no clear signs of progress toward creating such a plan.