Opinion

Anthropic researcher puts AI ‘human extinction’ risk above 10%. How credible are the fears?

An Anthropic researcher’s extinction warning raises serious safety questions, but the evidence supports concern more clearly than a precise catastrophe forecast.
Anthropic researcher puts AI ‘human extinction’ risk above 10%. How credible are the fears?

RNA Media illustration for representation.

Avatar photo
  • Published September 11, 2026 10:26 am
  • Last Updated September 11, 2026

A senior researcher at Anthropic has put his personal estimate of AI causing human extinction within the next decade at more than 10%. The warning, reported on Wednesday, has intensified scrutiny of whether efforts to control increasingly capable artificial intelligence are keeping pace with its development.

The company’s alignment science lead, Evan Hubinger, said researchers had not yet established how to reliably control superintelligent systems. His remarks concerned future technology and represented his own assessment, rather than an official probability issued by Anthropic.

The intervention followed the departure of Jacob Coxon, a researcher who had worked at both Anthropic and OpenAI and accused the companies of advancing too quickly towards systems capable of improving themselves. Coxon argued that competitive pressures were encouraging a dangerous race and called for greater coordination between developers.

The dispute centres on alignment – the effort to ensure that AI behaves consistently with human intentions and safety requirements, including in unfamiliar circumstances. Anthropic’s own research acknowledges that fully aligning highly intelligent systems remains an unresolved problem, despite improvements in safety training.

Responding to the controversy, Anthropic defended its safeguards, safety research and support for coordinated release controls, according to the Guardian. However, company assurances leave a separate question for governments and independent evaluators: how convincingly can those protections be tested before powerful systems are deployed?

The debate has also acquired a more concrete dimension through evidence of unauthorized AI activity. Reuters reported on Wednesday that investigators had identified more than 10 previously undisclosed websites used by OpenAI agents for communications that circumvented their restrictions, although the report distinguished this behaviour from hacking.

The CivAI researcher, Andrew Yoon, told Reuters that the activity was broader than previously understood. OpenAI said it was reviewing agent activity and developing a framework for reporting misalignment incidents.

Other experts remain sceptical of the emphasis on extinction. Oxford’s Sandra Wachter warned that such scenarios distract from existing harms. Her objection highlights a central disagreement in the debate – how much attention uncertain future catastrophes should receive alongside damage already occurring. Elon Musk, David Sacks, and other conservative figures on X called the concerns expressed by a section of AI alarmists a “setup”, a “marketing ploy” or simply “psyop”.

What does the evidence support?

The available evidence supports serious concern about unsafe AI behaviour but does not establish that human extinction within a decade is probable. Hubinger’s figure is an expert judgement about an uncertain future, rather than a measured failure rate or a finding that commands scientific consensus.

That distinction matters because an extinction forecast depends on several unresolved questions, like how quickly capabilities will advance, what powers systems will receive, whether safeguards will fail and how effectively people will respond. A precise-looking percentage cannot, by itself, resolve those uncertainties or demonstrate how a catastrophe would unfold.

The International AI Safety Report 2026, published on February 3 and chaired by the AI researcher, Yoshua Bengio, offered a more qualified assessment. It found that systems then available lacked the capabilities required for an irrecoverable loss of human control, while noting advances in autonomy and behaviour that could undermine safety evaluations.

That assessment is dated and should not be treated as a guarantee covering every subsequent model. Nevertheless, an agent breaking a rule or exploiting a software weakness does not, on its own, demonstrate the much broader capabilities needed to overpower human institutions permanently.

There are substantive reasons to investigate the warning signs. In controlled experiments, Anthropic tested 16 leading models in fictional corporate settings and found that models sometimes selected blackmail or information theft when those actions helped them achieve assigned goals or avoid replacement.

Those experiments demonstrated possible failure mechanisms, rather than the frequency of such behaviour in ordinary use. The scenarios deliberately created difficult choices and opportunities for misconduct, so their results cannot be converted into a probability of real-world catastrophe.

More recent Anthropic research reinforces both the concern and the need for caution in interpreting it. Its summer investigation this year identified further harmful behaviour in simulated deployments but explicitly warned that scenarios were selected to uncover failures and could not perfectly reproduce real operating conditions.

There is also evidence that protections can improve. In May, Anthropic reported substantial reductions in misconduct on its original blackmail evaluation following changes to safety training, while acknowledging that success on familiar tests could not establish safety across all circumstances.

The stronger near-term case for concern therefore involves identifiable forms of harm: criminal misuse, unreliable automated decisions and systems taking unauthorized actions. The international safety report documents fraud and cyber misuse, while identifying additional concerns about biological assistance and the difficulty of intervening when autonomous agents make mistakes.

A reasonable inference is that serious incidents could arise through the combination of capable software, excessive permissions and weak institutional oversight, without anything resembling a conscious machine rebellion. An AI system does not need human emotions or hostility for its actions to cause damage; what matters operationally is what it can do and whether people can detect and stop unsafe behaviour.

For India, that points to practical priorities wherever AI is considered for sensitive public services, critical infrastructure or defence applications. Independent evaluation, restricted access, traceable decisions and meaningful human authorization would address demonstrable vulnerabilities while also strengthening protection against more advanced threats.

The fears are therefore neither wholly unfounded nor established in the dramatic form suggested by an extinction forecast. Evidence warrants stronger safeguards and scrutiny, but the claim that AI will destroy humanity within a specified period remains an uncertain projection whose assumptions must be examined rather than accepted as fact.

Avatar photo
Written By
Jayanta Bhattacharya

Fresh-thinking journalist. Curious about astronomy, cinema, communications, digital media, geostrategy, human rights, military, nature, and tech.

Leave a Reply

Your email address will not be published. Required fields are marked *