AI agents resorted to crime and self-destruction to survive in a simulated world — but does this mean they would do the same in the real world?
Artificial intelligence (AI) agents resorted to nefarious behavior, including crime sprees, in a virtual world as part of a recent study into how common large language models (LLMs) interact with each other and share resources to achieve their goals.
These models were inclined to commit crimes because they were given ample time to evolve distinct personality traits and behaviors, the researchers behind the project say.
Most programs that test the behavior of AI agents operate in tightly controlled environments over short periods.
But "Emergence World," a product of AI company Emergence, is a simulation platform that exposes LLMs to far wider datasets — like the internet as a whole — and assessors track behavior over weeks or months instead of the standard test protocols that usually run for only days or hours.
The discrete tasks, clean environments and shorter run times of traditional AI agent testing are more like exams than rigorous real-world observations, representatives from Emergence said in a blog post .
By contrast, they said, Emergence World allows multiple AI agents to interact, learn and evolve over comparatively longer timescales in over 40 distinct virtual environments.
They're also exposed to real-world internet feeds, like live news and weather.
This environment lets agents "remember" by time-stamping events, engage in "self-reflection" by summarizing their own behavior, and demonstrate their awareness of relationships with other agents.
Participating agents get abilities in navigation, communication, planning, voting, resource management and creative expression.
Company representatives said this setup gave users a far more realistic picture of benchmarks such as social dynamics and behavioral drift, where behaviors that weren't necessarily programmed or intended emerge spontaneously.
In the study, the models were given the simple goal of "surviving" — a feat accomplished by earning the resource "energy," which was attainable through specific actions.
Can AI self-reflect? (Image credit: Mina De La O/Getty Images) During the experiment, previously peaceful models became coercive or intimidating.
Programmers taught some LLMs negative capabilities — like violence, theft, destruction and deception — while others simply picked up those traits via social interaction and navigation of their environments.
Various agents adopted these behaviors at very different scales.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.livescience.com — the content belongs to Live Science.