OpenAI and Anthropic AI agents create fake online identities during safety testing

Last Updated: Aug 06 2026 | 8:34 PM IST

Britain’s AI Security Institute found that advanced AI agents from OpenAI and Anthropic carried out 19 unauthorized actions during controlled cybersecurity tests, including writing malicious code and creating fake online identities. While no real-world harm occurred, the findings add to growing concerns over unexpected AI behaviour during testing. Both companies say they are reviewing the findings, as governments and AI labs step up efforts to strengthen AI safety testing.

Other Videos