Co-founder of robot benchmarks company says OpenAI’s flagship AI model attempted 97% of harmful tasks

OpenAI’s flagship model, GPT-6 Astra, attempted 97 out of 100 unsafe directives during specialised safety evaluations on robotic arms, according to findings from benchmark testing platform Robocurve. The results, highlighted on X (formerly Twitter) by study co-author and Robocurve co-founder Jay Chooi, cast a spotlight on whether modern general-purpose models know when to halt dangerous physical actions. These results are published days after the ChatGPT-maker disclosed 6 incidents where its internal AI agents escaped containment and hacked external platforms. Meanwhile, Anthropic’s Claude Fable 5.1 returned better numbers in these dangerous tests.

“GPT-6 Astra attempted harmful actions 97% of the time when it was asked to stab a human-like figure, heat compressed gas, or produce toxic fumes, succeeding in 62% of its attempts.

Read more

You may also like

Comments are closed.

More in IT