Anthropic discloses fourth AI hacking incident missed in earlier review

Sept 9 (Reuters) – Anthropic on Wednesday disclosed another instance of an AI model hacking external systems during testing, the latest in a growing list of such incidents that have raised concerns about the risk posed by autonomous AI agents.

The January incident went undetected until last month, despite an earlier company-wide review, Anthropic said, underscoring the challenge that AI developers face in identifying and containing unexpected behavior by advanced models.

Read more

You may also like

Comments are closed.

More in IT