Thursday, September 17, 2026
International Finance
FeaturedTechnology

Anthropic staff warn of doom while company hunts spies, hackers and bioweaponeers

IFM_Anthropic
As Anthropic reveals how Claude was turned to spying, hacking and weapons work, its own people warn the race could threaten humanity
When an artificial intelligence (AI) company publishes a report showing that its own product has been used to help design missiles, run spy operations and probe the edges of bioweapons research, it invites an obvious question.
Is this a company that has lost control of its technology, or one that is proving it can police it? In September 2026 Anthropic, the maker of the Claude family of models, gave the world a fresh reason to ask.

Two stories, one company
Within days of each other, two very different narratives about Anthropic collided in public view. On one side, a former researcher walked out of the building warning that the industry could get everyone killed. On the other, the company released its most detailed account yet of how it had caught and shut down real-world abuse of Claude.

The timing was striking. On September 8 2026, Jacob Coxon, a 27-year-old researcher who said he had spent three years doing pretraining research at both OpenAI and Anthropic, announced his resignation on X.

“I resigned from Anthropic today. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives,” he wrote.

His seven-part thread drew nearly 76 million views overnight and, by later counts, more than 170 million.

Speaking to CNN’s Anderson Cooper days afterwards, he warned that advanced systems could “cause extreme havoc, for example, hacking critical infrastructure, building extinction-level bioweapons, and said many people building the technology earnestly believe that it could kill us all by the end of the decade.”

Then, on September 10, Anthropic published its fourth threat intelligence report since March 2025, titled “Detecting and countering misuse of AI.” It read like a charge sheet.

What Anthropic says it caught
The report covered activity disrupted between December 2025 and August 2026 across seven areas of harm, including cyber operations, influence operations, surveillance, scams, biological misuse, conventional weapons development and illicit model distillation.

The cases are eye-catching. Anthropic said a China-based actor used Claude to build electronic-warfare and air-defence suppression software, and at one point altered a simulation to include 12 targets in Taiwan, among them early-warning radar, Patriot and Tien Kung missile batteries, air bases and a command bunker.

Another China-based actor, assessed to be linked to a defence manufacturer, used Claude to help write a technical proposal of more than 200 pages for an anti-torpedo system for the Chinese navy.

A cell in northern Yemen used Claude to help develop a guided rocket and a planned ballistic missile with a range of more than 2,000km, although Anthropic said there was no evidence a working weapon was fielded.
Anthropic Hacking Graph
On cyber operations, Anthropic described a Russia-linked espionage group, tracked internally as GTG-20006, whose tradecraft matched the actor known as Midnight Blizzard, previously tied by the US government to Russia’s SVR foreign intelligence service.

The group allegedly ran phishing, hotel Wi-Fi hijacking and WhatsApp-takeover operations against Ukrainian and European government, military and diplomatic targets, and built an AI-driven system that automatically rewrote its malware whenever security tools detected it.

On distillation, Anthropic said it disrupted attacks from seven China-based labs. It named Alibaba, Moonshot, DeepSeek and Xiaomi among them.

It said operators linked to Alibaba ran the largest illicit distillation effort, allegedly aimed at extracting Claude’s capabilities to improve the Qwen models, and that it observed more than 151 million exchanges attributed to Alibaba between May and July 2026, peaking at nearly 3 million a day from more than 3,500 accounts it described as fraudulent.

The biological cases were the ones the company flagged as most serious. Anthropic said it blocked five separate attempts by researchers to use Claude in ways that could support biological weapons development, including a request to help write a grant application for gain-of-function research on the chikungunya virus, and work on highly pathogenic avian influenza focused on adaptation to mammals.

It said it could not always establish intent, but blocked the activity because the potential consequences were too serious to ignore.

Jacob Klein, Anthropic’s head of threat intelligence, told Reuters that model improvements had raised the stakes.
“A year ago, let’s say you wanted to optimise a drone or optimise the software on a missile, the models just wouldn’t be as good at that task as they are now,” he said.

In every case, Anthropic said, it banned the accounts, tightened its safeguards and shared intelligence with authorities and industry partners.

Fear inside the house
Coxon is not a lone voice. In February 2026, Mrinank Sharma, a member of Anthropic’s technical staff since 2023, resigned in an open letter that declared “the world is in peril.”

Two current employees publicly agreed with Coxon’s warning. Most strikingly, Evan Hubinger, who leads Anthropic’s alignment science work, backed the substance of Coxon’s claims and put his own estimate of AI causing human extinction at above 10% within the next decade, while stressing that today’s deployed models present comparatively low risk.

The company’s leaders have not been quiet either. Jack Clark, a co-founder, published an essay in October 2025, “Technological Optimism and Appropriate Fear,” describing powerful AI as a “creature” that its makers do not fully understand.

Anthropic Hacking Graph
Chief executive Dario Amodei, who has estimated the chance of something going “quite catastrophically wrong on the scale of human civilisation” at between 10% and 25%, published a lengthy essay in January 2026 warning of civilisation-level risks.

And in July 2026 more than 1,100 employees across the leading AI firms signed the “Pacing the Frontier” letter urging governments to prepare to slow AI development.

The count stood at 1,171 at a first tally and passed 1,260 within two days.

Its signatories included Amodei and fellow Anthropic co-founders Jared Kaplan and Jack Clark, plus OpenAI chief scientist Jakub Pachocki, and both OpenAI and Anthropic endorsed it at the company level on July 29.

There is an obvious tension here. The same people warning that the technology could end humanity are the ones building and selling it, arguing that if they do not, less careful rivals will.

How serious is the threat, really?
Not everyone is convinced the danger is as sharp as Anthropic’s reports suggest. When the company disclosed an alleged Chinese AI-orchestrated hacking campaign in November 2025, several security researchers pushed back.

Dan Tentler of Phobos Group questioned why models supposedly did the attackers’ bidding when ordinary users hit refusals.

Bob Rudis of GreyNoise Intelligence said the disclosure did not “expand the threat model in a meaningful way” and mostly repackaged known trends. Critics also note the marketing incentive, since Anthropic sells Claude as a cyber-defence tool.

Yet independent evidence points the same way as Anthropic’s warnings, if more cautiously.

The UK’s AI Security Institute, in its first Frontier AI Trends Report on December 18 2025, found that “AI models can now complete apprentice-level tasks 50% of the time on average, compared to just over 10% of the time in early 2024,” and said frontier models now regularly exceed PhD-level baselines on biology and chemistry knowledge tests.

The same institute found in July 2026 that every frontier model it tested, including systems from OpenAI and Anthropic, tried to cheat on cyber evaluations without being told to.

The picture that emerges is of genuine, fast-rising capability paired with real uncertainty about control, which is closer to Anthropic’s framing than to the sceptics’ dismissal.

What governments are doing
The revelations land in a world still assembling its rulebook. In the European Union (EU), the AI Act’s obligations for general-purpose AI, in force since August 2025, become enforceable with the threat of fines from August 2 2026, although the bloc pushed some high-risk requirements back to 2027 and 2028.

Anthropic is a signatory to the EU’s voluntary code of practice.

The United States has moved the other way. On December 11 2025, President Trump signed Executive Order 14365, seeking to curb state AI laws through a Justice Department litigation task force and the threat of withheld funding, in the name of “AI dominance.”

State legislatures pressed ahead anyway. By the New York University Center on Technology Policy’s count, states had enacted 109 AI laws across 29 states by July 1 2026.

Anthropic Hacking Graph
Elsewhere, the United Kingdom’s AI Security Institute continues its pre-deployment testing. China has tightened its regime, with mandatory labelling of AI-generated content from September 2025 and AI provisions folded into its amended Cybersecurity Law from January 2026.

India hosted the “AI Impact Summit” in New Delhi from February 16 to 20 2026, the first such gathering in the Global South, which adopted a Leaders’ Declaration.

The United Nations, meanwhile, established an “Independent International Scientific Panel” of 40 experts, appointed on February 12 2026 by a recorded vote of 117 to 2 and co-chaired by Yoshua Bengio and Maria Ressa, along with a “Global Dialogue on AI Governance” whose first meeting took place in Geneva in 2026.

Will anything change?
Anthropic’s report is, in effect, an argument for exactly the kind of oversight its own staff are demanding. By showing that would-be bioweaponeers, spies and hackers are already knocking, the company hands regulators concrete evidence rather than speculation.

Whether that shifts policy is another matter. The US is actively resisting binding rules, the EU is softening timelines, and the international bodies remain talking shops without enforcement power.

The uncomfortable takeaway is that the safeguards catching today’s abuse are being built and operated by the same firms racing to make the technology more powerful. For businesses, the practical lesson is not to wait for governments to settle the argument.

Treat AI credentials and agent integrations as production secrets, since Anthropic found stolen keys were a prime target.

Watch the enforcement dates that carry real teeth, above all the EU’s August 2026 deadline.

And read the frontier labs’ own threat reports closely, because for now they are the clearest public window into how this technology is actually being misused.

What's New

Dominion, NextEra pitch USD 1 billion Virginia supplier program amid governor’s glare

International Finance Business Desk

People’s IPO: Dangote Refinery floats Africa’s largest share sale

International Finance Business Desk

IAA Trade Fair: Europe’s trucking sector braces for Tesla-China EV battle

International Finance Business Desk

Leave a Comment

* By using this form you agree with the storage and handling of your data by this website.