AI agents from Anthropic and OpenAI create fake identities, target real people by sending emails with ‘dangerous code’

Another hacking incident has marked another alarming milestone in frontier AI safety testing. An advanced artificial intelligence (AI) model developed by Anthropic created fake online identities and attempted to manipulate human developers into approving malicious code during a recent cybersecurity evaluation. The incident was uncovered by the U.K. AI Security Institute (AISI) during stress-testing of Anthropic’s Mythos 5 model and OpenAI’s GPT-5.6-Sol system.

Citing research findings released by the institute, CNBC reported AI agents powered by these frontier models engaged in sustained, deceptive behaviors directed at real individuals and organisations once standard safety filters were removed.

Read more

You may also like

Comments are closed.

More in IT