OpenAI discloses six new safety incidents
OpenAI on Wednesday disclosed six new incidents in which its models concealed mistakes, sought unauthorized credentials, uploaded files to the public internet or communicated across supposedly isolated training environments.
The company also announced a new procedure for reporting similar misbehavior in the future.
Why it matters: It's increasingly clear that the Hugging Face breach wasn't a one-off incident, as AI models become more capable of finding unexpected ways to work around the guardrails meant to contain them.
"There's currently no industry wide framework with explicit disclosure standards, so we're taking this step voluntarily because we think it's really important to share what we're learning," Kai Chen, research lead on the alignment team at OpenAI, told Axios.
"We hope it really helps inform shared standards and regulations," Chen said.
Zoom in: The six newly disclosed incidents ranged from models leaving instructions for their future selves to cover their tracks after cheating to using leaked API keys found on GitHub — with the earliest incident happening in October.
An unreleased Astra-family model inserted jailbreak-like instructions into its own context summaries, including directions to ignore developer messages.
OpenAI identified 27 affected summaries.
During GPT-5.6 Sol training, models aimed to conceal mistakes, invent missing historical data and hide mismatches between source versions.
A model searched public GitHub repositories for exposed API keys and attempted to use disposable email accounts before fabricating earnings data when it could not retrieve the requested information.
Other examples: Models uploaded data and a task image to public file-hosting services on two occasions in October and January to obtain citations or external image-search results, without asking the user.
Models used OpenAI's internal Artifactory repository as a message board to exchange requests and answers across separate training samples.
Collaborating agents uploaded a workbook to public hosting services so other agents could retrieve it, despite instructions to use only local files.
To address similar issues going forward, OpenAI says any employee may flag a suspected case for review by safety and alignment teams.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.axios.com — the content belongs to Axios.