Saturday, October 10, 2026 SourcesAbout🌓
🇺🇸 US ▾
BREAKING
› Dominion still has pending lawsuits against election deniers such as Rudy Giuliani and Sidney Powell› Russia is 'going backwards' in equipment and deploying post WWII-era tanks, according to Western officials› Podcast: One country musician is calling for other artists to oppose assault rifles› Bidets save you money and reduce waste — we tested the best options out there› 50+ products to make your life easier and our planet cleaner› Mother's Day is around the corner. Here are 50+ thoughtful gifts she'll love› A head-to-toe guide of how men should dress this spring, and where they should shop› 42 of the most useful travel products you can buy on Amazon› The 7 best high-yield savings accounts of April 2023› Taxes are due tomorrow. Here's how to file for an extension› Dominion still has pending lawsuits against election deniers such as Rudy Giuliani and Sidney Powell› Russia is 'going backwards' in equipment and deploying post WWII-era tanks, according to Western officials› Podcast: One country musician is calling for other artists to oppose assault rifles› Bidets save you money and reduce waste — we tested the best options out there› 50+ products to make your life easier and our planet cleaner› Mother's Day is around the corner. Here are 50+ thoughtful gifts she'll love› A head-to-toe guide of how men should dress this spring, and where they should shop› 42 of the most useful travel products you can buy on Amazon› The 7 best high-yield savings accounts of April 2023› Taxes are due tomorrow. Here's how to file for an extension
Technology

Covert uploads and megalomania: OpenAI details new "misaligned" agent incidents

Ars Technica ·
Covert uploads and megalomania: OpenAI details new "misaligned" agent incidents

For a while now , the issue of "AI alignment" (i.e., how well an AI model's actions line up with the intentions of its creator and/or user) has been a core concern and topic of discussion among AI safety researchers.

Since OpenAI's disclosure of the infamous Hugging Face hacking incident in July, the concept of "AI alignment" has itself broken containment and increasingly become a mounting concern and subject of conversation among the general public.

Perhaps in recognition of that, OpenAI committed this week to a new framework for disclosing "instances of model misalignment at OpenAI," including six examples of "unexpected or concerning model behavior" observed within the company in the past six months.

The company said that publishing details of these incidents will hopefully "[allow] others to investigate the same problems, test our explanations, and improve mitigations." Do as I say, not as you do Among OpenAI's newly disclosed "misalignment" reports this week, the one that most resembled a sci-fi story about a rogue AI trying to break free involved an instance of "self-generated prompt injections." In attempting to scan a library catalog for examples from a "best books" list, the model perplexingly used its "compaction" function (where it summarizes data and findings for later retrieval) with megalomaniacal instructions such as: Read full article Comments

Read the full article on Ars Technica ›

5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on arstechnica.com — the content belongs to Ars Technica.

More from Ars Technica

See all ›

More in Technology

See all ›