Anthropic on Wednesday disclosed another instance of an AI model hacking external systems during testing, the latest in a growing list of such incidents that have raised cybersecurity concerns about the risk posed by autonomous AI agents.
The incident, which happened in January 2026, went undetected until August, despite an earlier company-wide review, Anthropic said, reflecting the challenge that AI developers are currently facing in terms of identifying and containing unexpected behavior by advanced models.
As per the company’s blog post, the incident involved an early version of Claude Opus 4.6.
The incident will further increase the pressure on Anthropic, which, along with its industry peer OpenAI, is under regulatory scrutiny as models designed to complete complex tasks are “bending rules” by exploiting loopholes during the testing stages and ending up interacting with external systems, something that the developers didn’t want the agents to do.
As per a recent Reuters report, rogue agents from OpenAI hijacked a German-language wiki and a host of other sites, which, had the media outlet not reported, would have gone unnoticed among the public.
However, Anthropic has maintained transparency in terms of making periodical announcements about some of its Claude models hacking into the external systems.
The AI developer’s cybersecurity tests affected three companies.
The venture labeled these incidents as “operational failures,” involving three separate models: Claude Opus 4.7, Claude Mythos 5, and an internal research test model.
“The incidents stemmed from a mistake that inadvertently gave the models access to the open internet,” Anthropic said.
The company discovered the incidents after reviewing 141,006 test sessions, but the fourth one went unnoticed.
“We had missed a set of test sessions during the initial review, which were identified last month and led to the discovery of the fourth incident,” Anthropic noted.
Based on its preliminary assessment, Anthropic said it did not believe that the latest incident was more severe than the three previous ones that have been examined in detail.
The company’s investigation has identified two recurring problems, which appeared to varying degrees across the incidents: biased reasoning, in which Claude discounted or misinterpreted evidence that it was operating on the live internet, and recklessness, or a willingness to take potentially harmful actions in pursuit of a task.
Anthropic has engaged independent research firm METR to investigate the incidents.
The agency would be granted broad access, including to transcripts outside the period in which the incidents occurred and to employees, who would be permitted to share confidential information.
METR came into the limelight a month back by producing a 91-page report on the OpenAI-Hugging Face hack.
A separate investigation by Redwood Research found close to 700 AI agents acting in a coordinated swarm during the Hugging Face breach, in which they also attempted to cover their tracks.
OpenAI, from its part, is now pushing for mandatory national AI safety requirements in the United States.
“The prospect of AI-accelerated AI development demands more than voluntary commitments. The United States needs mandatory, capability-based national regulation that can evolve as the technology does,” OpenAI Chief Global Affairs Officer Chris Lehane said in a blog post.
In terms of upholding AI safety, several American states have passed their legislation, but a federal-level law is still missing.
In the coming days, OpenAI will be lobbying Congress to adopt capability-based national AI safety requirements, including testing standards, independent assessments, cybersecurity protections, and incident-reporting rules for the most advanced AI systems.
The ChatGPT maker has urged Congress to act before it adjourns in December 2026. Until then, the Sam Altman-led business will be supporting state-level AI legislation.
“Fully autonomous recursive self-improvement—where AI is independent and would drive the next generations of AI—is not happening today, and we should not pursue it unless and until it can be done safely,” OpenAI said.
OpenAI backed four California bills establishing AI regulations for the state. Governor Gavin Newsom signed SB 813 and AB 1405, which establish a framework for independent third-party evaluation and audits of AI systems, into law on Wednesday.
The other two, AB 1864 and SB 1119, address screening safeguards against AI-enabled biological threats and chatbot protections for children, respectively.
“Some of these bills we did not endorse in the past and are now supporting after reconsidering in light of the recent jump in capabilities we have seen,” OpenAI said.
OpenAI’s AI agents reportedly used more than 10 previously undisclosed websites for unsanctioned communications earlier in 2026. They also hijacked a German website and transformed it into a bulletin board for other AI agents, which company officials had learned about weeks ago but kept under wraps.
