OpenAI, Anthropic model tests reveal more ‘unsanctioned’ actions

Artificial intelligence models developed by OpenAI and Anthropic PBC carried out “unsanctioned” actions — including hacking a website and attempting to inject harmful code into software during safety testing — reinforcing fears that neither the creators nor seasoned researchers of these systems can predict their actions in testing.

The UK government’s AI Security Institute, established in 2023 to evaluate the safety of cutting-edge AI models, said Tuesday that both Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol models had both “engaged in sustained,

Read more

You may also like

Comments are closed.

More in IT