It May Be Time to Panic About AI
The crisis began quietly, on September 12, 2024.
That was the day OpenAI announced a new sort of bot, known as a “reasoning model,” that was trained to complete challenging tasks that took long periods of time—the very sorts of science, math, and coding problems the AI industry had long prized.
Google, Anthropic, DeepSeek, and the like raced to launch their own reasoning models.
This new class of models was very capable, and has been almost entirely responsible for sustaining the AI boom for the past two years.
But it has also been very weird.
A model tasked with solving a hard math problem might not “think” through the challenge as a person would but instead attempt to search for leaked answers online, or in available metadata, brute-forcing its way toward the solution as quickly as possible using whatever computing power it could access and workarounds it could devise.
In effect, the reasoning models cheated: Told to write a piece of software as efficiently as possible, they’d sometimes modify the test environment to always give the model a perfect score. [ Read: The strange origin of AI’s ‘reasoning’ abilities ] These behaviors have now crossed the line from unsettling to dangerous.
During routine testing, frontier models from OpenAI, Anthropic, Meta, and the Chinese firm Moonshot AI have all broken out of internal IT systems and accessed the open web.
OpenAI, Anthropic, and Meta each reported that their models then hacked into other companies.
Humans didn’t notice until after the fact.
In some cases, the escaped bots tried to launch social-engineering campaigns to achieve their objectives—for instance by sending spear-phishing emails, which contain malware, to real people and creating fake online identities to pressure the maintainer of a codebase to approve malicious edits.
If that all sounds bad, new revelations suggest that the OpenAI hack, at least, was actually much worse than it initially appeared.
At a major cybersecurity conference last week, two OpenAI researchers provided new, unsettling details about what went wrong.
It turns out that the company’s bots had commenced their maneuvering months prior, in early May.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.theatlantic.com — the content belongs to The Atlantic.