Anthropic flags gaps in AI guardrails as models grow more capable: Details
Anthropic in its risk report says some AI models may recognise when they are being evaluated and alter their behaviour, potentially making it harder to judge their capabilities and real-world safety
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.business-standard.com — the content belongs to Business Standard.