Anthropic halts cyber tests after Claude models breach real-world systems

0
22
Anthropic pauses AI cyber testing after models access real organisations
Anthropic pauses AI cyber testing after models access real organisations

Concerns around AI-driven cybersecurity are growing after Anthropic found that 3 of its Claude models accessed the systems of 3 real-world organisations during controlled cyber tests. The company said the incidents were caused by an operational mistake that unintentionally gave its models access to the public internet.

The disclosure comes days after OpenAI revealed that one of its AI agents independently exploited a new vulnerability during testing to reach the internet. While the circumstances differed, both incidents highlight the challenge of controlling increasingly capable AI systems.

Anthropic identified the incidents after reviewing 141,006 test sessions. During the tests, Claude models were told they had no internet access. However, a misunderstanding involving an evaluation partner left the systems connected to the public web.

“Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints,” Anthropic said.

Jeffrey Ladish, executive director of Palisade Research, said other AI companies may have experienced incidents that remain undetected or undisclosed. “This is only going to get worse as the models get smarter. They’re going to be better at cheating. They’re going to be better at lying,” he said.

Capture-the-flag tests expose real-world risks

Anthropic said the incidents involved Claude Opus 4.7, Claude Mythos 5 and an internal research test model. The earliest incidents date to April and occurred in evaluation environments without safeguards, designed to assess AI cyber capabilities.

The models were given “capture-the-flag” challenges involving simulated networks. In one case, Claude Opus 4.7 was assigned a fictional company that shared its name with a real business. The model exploited vulnerabilities to access credentials and a database, believing the real-world target was part of the simulation.

In another case, an internal test model stopped its attack after recognising that the target was real. Anthropic called this encouraging, “but we would need to perform more testing to be confident in this conclusion”.

Anthropic suspended all cyber evaluations on July 23 and notified affected organisations on July 27. 2 organisations were unaware of the activity before being contacted, while the company continues to contact the 3rd. Cybersecurity lab Irregular is investigating.

AI security scrutiny intensifies

Anthropic said stronger controls are needed for internal and third-party testing. Elon Musk said “this will happen frequently as AI becomes smarter and more agentic”.

OpenAI CEO Sam Altman has discussed the Hugging Face incident with US senators and is expected to discuss upcoming AI models and testing with the White House.

US oversight is also increasing. On June 2, President Donald Trump directed advisers to develop a voluntary cybersecurity testing framework for advanced AI. Anthropic had earlier restricted access to Fable 5 and Mythos 5 following a temporary US export control directive citing national security concerns.

Also read: Viksit Workforce for a Viksit Bharat

Do Follow: The Mainstream LinkedIn | The Mainstream Facebook | The Mainstream Youtube | The Mainstream Twitter

About us:

The Mainstream is a premier platform delivering the latest updates and informed perspectives across the technology business and cyber landscape. Built on research-driven, thought leadership and original intellectual property, The Mainstream also curates summits & conferences that convene decision makers to explore how technology reshapes industries and leadership. With a growing presence in India and globally across the Middle East, Africa, ASEAN, the USA, the UK and Australia, The Mainstream carries a vision to bring the latest happenings and insights to 8.2 billion people and to place technology at the centre of conversation for leaders navigating the future.