Anthropic’s Claude Closes 85 Percent of AI Safety Gap in Automated Research Test
Anthropic researchers have tested Claude as an automated alignment researcher, allowing it to develop and evaluate methods for addressing 10 types of AI safety failures.
Claude closed up to 96 percent of the measured safety gap and achieved an average of 85 percent on deception tests.
The researchers also used the weaker Claude Sonnet 5 to improve alignment in an early Opus 4.8 model.
The results suggest AI could eventually handle more parts of AI research, although the experiment falls short of true self-improvement.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.gadgets360.com — the content belongs to Gadgets360 (NDTV).