AI can explain how it ‘thinks’ – but can we still trust it? OpenAI finds a problem

AI reasoning models are becoming better at showing how they arrive at an answer. Their chain of thought can also give researchers a window into whether a system is following instructions or behaving in ways that could create safety risks.

But OpenAI’s latest tests have exposed a problem. What if the AI becomes better at controlling what researchers can see?

OpenAI’s GPT-6 Astra system card says the model is significantly more capable than GPT-5.6 Sol at controlling its own chain of thought (CoT). Researchers also found that Astra can sometimes reduce or alter its visible reasoning when it knows it is being monitored.

Read more

You may also like

Comments are closed.

More in IT