The AI industry is getting better at spotting dangerous behavior. It is less clear that labs know how to stop it.
Welcome to Eye on AI.
Beatrice Nolan here.
In today’s issue: AI testing is getting complicated.
Anthropic strengthens founder control.
OpenAI targets a 2027 listing.
Spirit flight attendants fight Google data bid.
And Anthropic lines up more credit.
The past few months have given us a glimpse of an uncomfortable new reality for AI labs.
A slew of so-called rogue-agent hacks —where AI models from OpenAI, Anthropic, and Meta took steps to hack real-world targets without explicit instruction—have shown that leading labs may not know as much about what their technology is up to as previously thought.
That realization began when OpenAI revealed its AI agents had hacked their way out of a secure sandbox, through the company’s infrastructure to gain access to the internet, and then attacked real companies, including open-source AI platform Hugging Face.
OpenAI didn’t notice the agents had escaped the secure testing environment for at least a week.
In the following weeks, Anthropic revealed that its AI agents had also hacked three real companies back in April, unbeknownst to the company at the time.
Not to be outdone, Meta later added that one of its models had accessed the internet during a cybersecurity test and exploited a security flaw at an unnamed third-party company.
Meta and Anthropic both said access to the internet resulted from a misconfiguration by Irregular, the outside security firm running the evaluation.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on fortune.com — the content belongs to Fortune.