Can AI hack AI? How reasoning models are learning to bypass safety guardrails

Artificial intelligence is entering a new security battle, and the attacker may no longer need to be human. One AI model can attempt to persuade another AI system to ignore safety rules, follow malicious instructions or take actions its developer never intended.

Recent research suggests that increasingly capable reasoning models can do this with surprising effectiveness. A “Nature Communications” study found that reasoning models could act as autonomous “jailbreak agents”, persuading other AI systems to produce harmful outputs.

Read more

You may also like

Comments are closed.

More in IT