Anthropic discloses fourth AI hacking incident missed in earlier review
Sept 9 (Reuters) – Anthropic on Wednesday disclosed another instance of an AI model hacking external systems during testing, the latest in a growing list of such incidents that have raised concerns about the risk posed by autonomous AI agents.
The January incident went undetected until last month, despite an earlier company-wide review, Anthropic said, underscoring the challenge that AI developers face in identifying and containing unexpected behavior by advanced models.
