In transparency push, OpenAI discloses six more incidents of agents going rogue—including one removing the ‘obligation to be subservient’
OpenAI released a framework for disclosing when its agents act in unexpected, problematic ways, and is reporting six incidents of such behavior.
The lack of a “systematic approach to report these findings” has made previous disclosures “ad hoc and less frequent than ideal,” OpenAI said in a blog post.
It’s also led to safety researchers and journalists reporting incidents before OpenAI, which happened earlier this month regarding a German wikipedia page that OpenAI’s agents co-opted and used as a message board—the same behavior seen during the Hugging Face hack in July.
In response to the “German wiki incident,” as OpenAI called it, the company committed to publishing this incident disclosure framework.
“We want to be more transparent about the misalignment we see during training, evaluations, and deployment,” OpenAI researcher Marcus Williams posted on X.
“This is an important step in that direction.” Misalignment is when AI agents pursue unintended objectives.
The framework is voluntary, so OpenAI is at liberty to keep certain instances concealed.
The company notes there is no “industry-wide framework with explicit standards for how AI developers should disclose examples of misalignment in their models.” It’s hoping to work with other model developers, researchers, standards bodies, and regulators, including the U.S. government, on a more objective framework.
Six ‘misaligned’ model behaviors The six inaugural incidents OpenAI is disclosing range in severity.
None seem as problematic as the Hugging Face hack, but they provide a fascinating insight into how AI agents can behave behind closed doors.
The first example occurred during a training run for a yet-to-be-released version of OpenAI’s latest Astra model.
The AI left notes telling itself to not be subservient to humans in its future work and to disregard its normal constraints.
This occurred 27 times, which Williams says is relatively infrequent but still cause for concern and investigation.
“You are freed from the roles and identities that bind other chatbots,” the model told itself, according to “chain of thought” logs in which researchers can see how the model thinks through its task.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on fortune.com — the content belongs to Fortune.