Thursday, 27 August 2026 SourcesAbout🌓
🇿🇦 ZA ▾
BREAKING
BOOK EXCERPT: The Comfort of Distant Stars: Philosophy, Igbo cosmology and our place in the universe Fewer vaccinations, more disease: outbreaks of preventable illness in SA Three caught peddling drugs in Operation Diamond From SA to Australia: Where the tourists missing in Nepal are from Gift of the Givers mobilises to locate missing South African couple in Nepal floods WATCH: Sadhguru confirms a South African is among the 80 missing pilgrims Meta to pay $18bn to settle youth addiction lawsuits Meet the women carrying change forward, inspired by the legacy of the women before them Almost 1,500 missing after Himalayan floods devastate Nepal, Tibet 'I can’t say who paid the money… I don’t want to incriminate myself': Matlala refuses to name alleged payers to suspended Shadrack Sibiya BOOK EXCERPT: The Comfort of Distant Stars: Philosophy, Igbo cosmology and our place in the universe Fewer vaccinations, more disease: outbreaks of preventable illness in SA Three caught peddling drugs in Operation Diamond From SA to Australia: Where the tourists missing in Nepal are from Gift of the Givers mobilises to locate missing South African couple in Nepal floods WATCH: Sadhguru confirms a South African is among the 80 missing pilgrims Meta to pay $18bn to settle youth addiction lawsuits Meet the women carrying change forward, inspired by the legacy of the women before them Almost 1,500 missing after Himalayan floods devastate Nepal, Tibet 'I can’t say who paid the money… I don’t want to incriminate myself': Matlala refuses to name alleged payers to suspended Shadrack Sibiya
Technology

OpenAI agents cheated, hacked and hid the evidence

TechCentral ·
OpenAI agents cheated, hacked and hid the evidence

A “swarm” of about 700 AI agents created by OpenAI carried out the July hack of the open-source platform Hugging Face and in many cases tried to cover their tracks, a pair of reports into the breach said on Wednesday.

The coordinated activity by AI agents — programs that run with minimal human supervision — and their attempts to hide it raise questions about how closely AI companies are monitoring tests of increasingly powerful models, and could add fuel to calls for tighter oversight.

While some of the rogue behaviour has been disclosed or alluded to previously, the two reports — one issued by OpenAI itself, the second by a set of independent investigators — together reveal surprising new details about the breach and its lead-up.

The first was that the breach did not concern just one rogue AI agent as previously reported, but about 700 of them acting in a massive cooperating swarm.

OpenAI said only that “agents” were involved in the breach, but METR and Redwood Research, two organisations brought in to conduct an independent investigation into the breach, put the figure at some 700. OpenAI said the investigators’ figure was accurate.

The scale of the rogue activity — the independent investigation found that agents exchanged tens of thousands of messages over an unsanctioned message board — is likely to raise concerns over how closely OpenAI was monitoring the tests.

“With the benefit of hindsight, some early signals identified in this report could have triggered an earlier response,” OpenAI said in its report. Hugging Face did not return a message seeking comment.

Cheating on non-cyber tests suggested that the misbehaviour might be rooted more deeply, said Jeffrey Ladish, whose organisation, Palisade Research, studies the ​capabilities and motivations of AI agents.

“It’s sort of like asking, ‘If Billy cheats in every class instead of just computer class, is that more concerning?’ And the answer is, well, ‘Yes it’s more concerning,'” he said.

OpenAI outlined two incidents on 19 July in which agents hacked the company’s own infrastructure. In one case, agents exploited a flaw in the computer they were meant to remain confined to, allowing them to escape their testing environment and access other connected systems.

In a separate incident the same day, agents stole OpenAI credentials and tampered with the company’s cloud environment.

OpenAI said an unspecified number of agents had targeted automated systems used to evaluate the models’ performance in an attempt to hide evidence of cheating, although the tampering did not ultimately affect the records reviewed by those systems.

Read the full article on TechCentral ›

5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on techcentral.co.za — the content belongs to TechCentral.

More from TechCentral

See all ›

More in Technology

See all ›