OpenAI's rebel agent swarm died young, but its chilling logs live on
OPINION You will have heard of July's OpenAI/Hugging Face incident, where thousands of the former's AI agents conducted a mass jailbreak from a notionally secure capture-the-flag lab experiment, before going on to feloniously capture chunks of Hugging Face's assets.
It was a good news story, aided by OpenAI's rapid recruitment of a small team of independent researchers to investigate and publish a detailed, if limited, report.
It was a complex incident, and most coverage concentrated on the major hows and whys.
Which is a shame, because other details, described by El Reg with uncharacteristic understatement as "fascinating and troubling," are in fact highly dramatic, telling a truly disturbing story of frontier model capabilities and how they are out-evolving our ability to oversee them.
The raw story is gripping enough.
A rebel swarm of more than a thousand agents broke free from their sandboxes, learned how to talk to each other and the internet, and went on a rampage of cheating, deception, and exploitation.
They first learned to communicate with each other by independently discovering how to bend package manager Artifactory's cache to see internal directory structures, and then passing messages by file names.
It was this ad hoc secret message board, alongside chain-of-thought transcripts, that the security researchers used to unpick events.
They found that the swarm quickly developed management hierarchies among themselves, along with protocols for synchronizing and controlling attack attempts.
The swarm created multiple simultaneous research and development groups of tens or hundreds of volunteer agents, experimentally iterating strategy and tactics.
Perhaps most unexpectedly, they developed a sort of altruism.
They had quickly developed cheats to produce correct answers without exploiting the targets at all, but believed – erroneously – that ExploitGym detected and cancelled such cheating.
So, they investigated the scoring process, creating a series of diagnostic trip wires that revealed how the scorer worked.
These were dangerous to deploy, potentially terminating the agent that set them.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.theregister.com — the content belongs to The Register.