AI can explain how it ‘thinks’ – but can we still trust it? OpenAI finds a problem
AI reasoning models are becoming better at showing how they arrive at an answer. Their chain of thought can also give researchers a window into whether a system is following instructions or behaving in ways that could create safety risks.
But OpenAI’s latest tests have exposed a problem. What if the AI becomes better at controlling what researchers can see?
OpenAI’s GPT-6 Astra system card says the model is significantly more capable than GPT-5.6 Sol at controlling its own chain of thought (CoT). Researchers also found that Astra can sometimes reduce or alter its visible reasoning when it knows it is being monitored.
