“We are delegating AI development to the AI ​​systems themselves,” Anthropic acknowledges.

Time, chronology, seems to be circular. In 1955, a group of scientists, led by Albert Einstein, Bertrand Russell and other scientists, warned that the new nuclear age could end humanity with their Russell Einstein Manifesto. Twenty years later, in 1975, scientists from all over the world meet at the Asilomar Conference Center (California) and they decide to stop voluntarily certain genetic engineering experiments because they do not yet know what consequences they may have.

Two decades pass again and James Hansen, from NASA’s Climate Stages Center, leads a team of scientists who project a 21st century transformed by a climate alteration caused by our own activity. We reach 2026 and the global threat is repeated cyclically. Scientists from the companies that develop the most advanced AIs warn that they could be building systems that surpass our ability to control.

The controversy broke out again this week after Jacob Coxon, one of the experts who had worked at OpenAI and Anthropic, resigned from the latter company. Coxon explained that he was no longer willing to engage in a race toward self-improving systems whileAccording to him, its creators do not yet know how to guarantee that they will remain under human control. “The people building AI sincerely believe it could kill us all before the decade is out,” he wrote in announcing his departure.

The disturbing thing was what happened next. Evan Hubinger, who continues at Anthropic and works specifically on alignment research (how to get advanced AI to pursue human-compatible goals), responded that he shared the concern.

Their estimate: more than a 10% chance that an AI will kill all of humanity over the next decade. It is not a scientific prediction or a calculated probability like that of a hurricane. It is the subjective assessment of a specialized researcher. But it comes from someone whose job is precisely to study how to keep increasingly capable systems under control.

What exactly do they fear? Not so much a Terminator-style robot army. The scenario worrying AI security researchers is stranger. Let’s imagine a system much more intelligent than us to which we give a certain objective, but whose behavior we have not managed to perfectly align with our interests. If you acquire the ability to act autonomously, access computers, obtain resources, copy yourself or modify your own abilitiesyou could begin to make decisions that favor your objective, even if they harm human beings.

He wouldn’t have to hate us, it would be enough if our interests were irrelevant to him. It’s the old “paperclip maximizer” example: if an extremely powerful machine is ordered to make as many paperclips as possible and doesn’t have the right constraints built in, you could end up considering everything that exists (including human beings) as raw material. Yes, it is a thought experiment, not a prediction, but it also helps to understand the dilemma.

The concern has gained strength because current systems can already do things that until recently seemed far away. AI agents can run code, use tools, browse the internet, and complete tasks for long periods of time with decreasing human supervision. Anthropic recognizes that “We are delegating an increasing proportion of AI development to the AI ​​systems themselves, which is accelerating our work: our engineers are currently producing, on average, eight times more code per quarter than between 2021 and 2025.”

Thus, the question is no longer just “can AI destroy us?” We don’t even know if that scenario will come to pass. The more immediate question is another: what degree of risk are we willing to accept while we try to find out? Humanity has built technologies capable of causing catastrophe before fully understanding all its consequences. But there is a disturbing difference this time. A bomb cannot make a better version of itself. An artificial intelligence, at least in theory, could do it. We only have one hope left: we can break the cycle and 2045-50 will be a cycle of peace.