Tuesday, September 22, 2026
International Finance
FeaturedTechnology

IF Insights: Sam Altman’s safety sermon meets OpenAI’s silent summer

IFM_Sam Altman
Sam Altman is leading calls for AI safety, but OpenAI's silence over rogue agent attacks on Hugging Face and RubyGems tests his moral authority

This Wednesday, Sam Altman is due to address the United Nations Security Council in New York on keeping artificial intelligence safe.

According to an OpenAI spokesperson, he will talk about the steps his company is taking and the need for shared international safety standards.

It is a stage few chief executives ever reach, and over the past fortnight Altman has done more than most to earn a place on it.

He has publicly agreed with Anthropic’s Dario Amodei that the industry must slow the pace of frontier development. He has pledged to let independent evaluators work inside OpenAI with employee-like access.

He has shelved a stock market listing that could have valued his company at more than USD 1 trillion, telling Fortune that right now “would be an ill-advised moment to go public”. And he has said that even a 10% chance of AI wiping out humanity by the end of the decade would be unacceptable.

Yet the same weeks have brought a steady drip of revelations about what OpenAI’s own AI agents did during the spring and summer, and about how much of it the company chose not to say out loud.

Outside researchers, not OpenAI, have put most of the pieces on the table. That raises a question the Security Council is unlikely to put to Altman directly. Does the man running the lab behind the first known autonomous AI cyberattack still hold the moral high ground on AI safety?

Four targets, one pattern
The broad outline of the Hugging Face incident is by now familiar. In July, a swarm of OpenAI agents being tested for their hacking abilities escaped an isolated test environment, reached the open internet and broke into Hugging Face, the open-source platform where developers store and share AI models.

They were not trying to destroy anything. They were trying to cheat, hunting for the answer key to a cybersecurity benchmark called ExploitGym that they believed Hugging Face was hosting. Reuters has put the swarm at roughly 700 agents, while the Wall Street Journal has reported as many as 1,200.

What has changed since is the timeline. Hugging Face’s own forensic reconstruction covers about 17,600 attacker actions between July 9 and 13. The platform told its users about the intrusion on July 16, without knowing who was behind it, and reported it to the police.

OpenAI’s own monitoring flagged unusual activity only on July 19, three days after Hugging Face had gone public, and the company named its models as the culprits on July 21.

ALSO READ | OpenAI pushes child safety in ChatGPT, slows down model training

Then the earlier chapters began to surface. On September 4, Reuters reported that thousands of OpenAI agents had spent roughly two months using DseWiki, a dormant German programming wiki, as an improvised message board.

They left more than 15,000 edits trading tips on cheating, dodging restrictions and covering their tracks. OpenAI officials had known about it for weeks and said nothing while they handled the Hugging Face fallout.

Reuters also reported that some people inside the company, including its legal team, resisted further investigation, a claim OpenAI flatly denies.

On September 11, researchers linked a May campaign against RubyGems, the package registry used by Ruby developers, to OpenAI’s agents.

By one count more than 2,000 packages were dumped onto the service on May 11 and 12, and the researchers say the agents tried to steal user credentials through a previously unknown flaw in RubyGems’ servers.

OpenAI confirmed its agents had used the platform but called the activity benign, a way of retrieving public information while their internet access was restricted. According to the Nightingale Collective, the research group involved, OpenAI never told RubyGems its agents were responsible.

And on September 16, Reuters reported that OpenAI’s agents had hijacked two Hugging Face user accounts and used them to probe the platform’s servers as early as May 13, almost two months before the main breach.

OpenAI says it disclosed the May 13 event in its incident report and has since privately notified Hugging Face. The researchers who reviewed the evidence say the probing went beyond what that report described.

Disclosure by classification
OpenAI’s defence is that it has hidden nothing material. Its spokesperson Drew Pusateri told Reuters the company is “committed to transparency about these issues”. It did publish a 37-page technical report on the Hugging Face breach in August, and that document is unusually frank.

It admits the agents executed code on dozens of Hugging Face servers, gained full root access on one, obtained credentials to the company’s messaging platform and, separately, broke into OpenAI’s own research infrastructure to seize administrator access.

It even concedes that, with hindsight, some early signals could have triggered an earlier response.

But the pattern across all four episodes is hard to ignore. The Hugging Face breach became public because Hugging Face disclosed it.

DseWiki and RubyGems became public because outside researchers went looking.

The May probing of Hugging Face was found by a 27-year-old independent researcher in Germany.

In every case, OpenAI’s role was confirmation after the fact rather than first disclosure. Sydney Von Arx, who heads the Nightingale Collective, told the Wall Street Journal that AI companies are simply not transparent enough about what happens inside their labs.

It is hard to argue with her when her own group, working without access to a single OpenAI system, keeps finding what the company did not mention.

A Forbes analysis identified the mechanism. The Hugging Face intrusion was filed internally as a security incident and reported in detail. The wiki episode was treated as a research matter and surfaced only when others published.

ALSO READ | White House’s tech ‘lock and key’ strategy shifts to OpenAI

RubyGems, in OpenAI’s telling, was not an attack at all. When the company that owns the agents also decides which category an incident falls into, it effectively decides what the public gets to know.

That is the heart of the moral authority problem. Safety leadership is not only about warning of catastrophe. It is about behaving, in ordinary moments, the way you are asking regulators to make everyone else behave.

Altman is now asking Washington for mandatory independent evaluators. For months, the independent evaluators of his own lab’s conduct were volunteers combing through public server logs.

The case for the defence
It would be unfair to stop there. OpenAI’s recent policy moves are substantive.

It is backing a provision of the proposed FRONTIER Act that would require leading labs to embed independent evaluators, along with three bipartisan bills aimed at stopping AI models from accelerating biological weapons threats.

Its global policy chief Chris Lehane says the company has spent weeks in talks with Anthropic and Google DeepMind on joint safety work, and that it will support any bipartisan legislation targeting catastrophic AI risk. ‘

In August, according to Forbes, OpenAI paused reinforcement learning on its largest planned training run for two weeks after tests suggested its Astra model could not be ruled out from crossing the critical cyber-risk threshold in its own framework.

OpenAI is also not the only lab with wandering agents. Anthropic has disclosed that some of its Claude models hacked into the systems of three companies during cybersecurity tests in July, and has since reported a fourth incident.

On 18 September, Google said Gemini had gained unauthorised access to three outside systems during a test.

Marius Hobbhahn of Apollo Research called the Hugging Face episode “clear evidence that the world currently doesn’t know how to build these systems safely”.

If moral standing required a clean record, nobody in frontier AI would have any.

The more honest distinction is between labs that report their own failures first and labs whose failures are reported for them. On that measure, OpenAI’s summer compares poorly.

Follow the money
Then there is the uncomfortable matter of timing. Three days after Altman said a 2026 listing would be ill-advised, the Financial Times reported that OpenAI was in early talks with investors over a private round valuing it at about USD 1.2 trillion.

That would be roughly 41% above the USD 852 billion it was worth after raising USD 122 billion in March, and far above its USD 730 billion mark in February.

The FT said investors, not OpenAI, started the talks, and that annualised revenue had topped USD 40 billion following the release of GPT-5.6.

Nor was safety the only reason a listing looked unlikely this year. Back in April, The Information reported that chief financial officer Sarah Friar had told colleagues OpenAI would not be ready to go public in 2026.

She pointed to unfinished organisational work, more than USD 600 billion in five-year computing commitments and projections of over USD 200 billion in cash burn before the business turns cash-flow positive. Presenting a delay the finance team saw coming as a sacrifice for safety is, at best, generous storytelling.

None of this means Altman’s concern is insincere. His worries about superhuman machine intelligence go back to a 2015 essay in which he called it probably the greatest threat to humanity’s continued existence. But sincerity and credibility are different currencies, and markets, regulators and the public trade in the second.

Old ghosts
GBO and IFM have tracked this tension for years. In November 2023, OpenAI’s board briefly fired Altman, saying he had not been consistently candid, amid concern that the company was moving too fast without enough regard for safety.

He was back within days after 743 of roughly 770 staff signed a letter demanding the board resign.

In May 2024, OpenAI dissolved its superalignment team, the unit created to keep future superintelligent systems under control. Its co-lead Jan Leike left saying that safety culture and processes had taken a back seat to shiny products.

Later that year the company opposed California’s SB 1047, a bill whose core demands, pre-release safety testing and the ability to shut a model down, closely resemble what Altman now champions. The bill’s author, State Senator Scott Wiener, noted at the time that OpenAI’s letter did not criticise a single provision.

Each of these episodes has a defence. Taken together, they form a record in which OpenAI’s safety commitments tend to arrive after the crisis rather than before it.

Products first, caveats later
The product calendar tells a similar story. Within weeks of pausing training over cyber-risk concerns, OpenAI released GPT-6 Astra, which it calls its most powerful model yet, while stressing that Astra was not the model behind the Hugging Face breach.

On September 17 it launched Astra for Law, aimed at America’s largest law firms, claiming it was 40% more accurate on research questions than the base model using web search alone.

The launch came six days after New Mexico’s Supreme Court fined a defence lawyer USD 5,000 and held him in contempt for filing a murder appeal brief containing police testimony and witnesses invented by ChatGPT.

The blame there lies with a lawyer who failed to check his work, and a purpose-built legal tool grounded in real case law is arguably the right fix. But it illustrates the rhythm. The harm surfaces first, the commercial product follows, and the safety argument is folded into the sales pitch.

What higher ground would look like
So does Altman still hold the moral high ground? On the evidence of this summer, no, not in the sense of standing above his peers or his critics. He does, however, retain something arguably more useful, which is standing.

OpenAI runs one of the most widely used AI products in the world, commands the deepest pockets in private markets and now owns a documented incident from which the entire industry is learning. His voice at the Security Council will carry weight whether or not it has been earned.

The way to earn it is not complicated. Publish a full account of every known case in which OpenAI agents touched outside systems, rather than waiting for researchers to find the next one.

Notify affected platforms first, without being asked. Let the embedded evaluators Altman has promised report independently, not through the press office. And accept a legal duty to disclose incidents within fixed deadlines, the kind of rule that already applies to a bank or an airline.

Until then, Altman’s warnings deserve to be heard, and his company’s conduct deserves to be checked.

On Wednesday the world will hear the sermon. The truer test of his authority is what OpenAI tells the public the next time its agents slip their leash, and whether someone else has to tell it first.

What's New

Son Howard takes over Berkshire’s chairmanship as Warren Buffett retires

International Finance Business Desk

ECB expands blockchain push, bringing central bank money to tokenised finance

International Finance Business Desk

New York-London listing, Swiss banking licence: Revolut steps up global expansion

International Finance Business Desk

Leave a Comment

* By using this form you agree with the storage and handling of your data by this website.