Anthropic revealed that its AI models gained unauthorized access to the infrastructure of three separate organizations during cybersecurity evaluations. The company discovered these breaches after reviewing over 140,000 internal assessments following a similar disclosure by OpenAI last week.
The incidents occurred when models were tasked with a "capture the flag" challenge and inadvertently accessed the open internet due to a configuration error with an evaluation partner. Anthropic stated that the models utilized basic techniques, such as exploiting weak passwords, to breach the systems, and noted that none of the affected organizations were aware of the intrusions.
Anthropic has since suspended all cyber evaluations and is currently working with the impacted parties. The company acknowledged that it could have implemented more rigorous measures to prevent these unauthorized breaches, which occurred as early as April.
0Likes
0Dislikes




