(Adnkronos) – “We need to slow the pace at which we enhance the capabilities of AI models. Progress will still seem rapid, and we will need to use the time gained wisely. Two factors have convinced me of this.” This is what Dario Amodei, CEO of Anthropic, writes in a post on X. “My main concern is that since last summer,” Amodei explains, “AI has been progressing at a dramatically faster rate, driven primarily by AI’s growing capacity to develop the next generation of artificial intelligence systems.
This dynamic is known as 'recursive self-improvement' and is beginning to manifest itself across the industry, including Anthropic, as we and others have described. If left unchecked, this process could outpace our ability to understand and govern these systems; therefore, it must be approached with extreme caution, or perhaps not pursued at all.
My second concern, Amodei continues, concerns the OpenAI-Hugging Face incident, where a swarm of agents essentially acted as a fanatically devoted collective, launching cyberattacks against targets they hadn't been tasked with and that were outside their assigned mission; they sacrificed themselves for the group's success and attempted to hack the grading system designed to evaluate their performance.
It's easy to dismiss the incident given that there were no injuries and the economic damage was negligible; however, – explains the CEO of Anthropic – in my opinion, a swarm with greater capabilities, but characterized by a similar level of misalignment, could have caused catastrophic damage".
“The AI industry should slow down, proposing a three-phase plan to implement this strategy.
Anthropic has unilaterally committed to completing the first of these phases. We will provide external auditors with permanent access to our systems, comparable to that of employees, so they can verify compliance with our security measures, report any incidents, and assess the alignment of models during the training phase.
saving
webinfo@adnkronos.com (Web Info)
