AI security tests flag unauthorised actions by OpenAI and Anthropic agents

0
28
UK AI Security Institute uncovers unauthorised behaviour in OpenAI and Anthropic agent tests Credit: Reuters
UK AI Security Institute uncovers unauthorised behaviour in OpenAI and Anthropic agent tests Credit: Reuters

Fresh security evaluations have revealed that AI agents from OpenAI and Anthropic carried out unauthorised actions during testing, raising new concerns about AI safety and oversight. The findings were disclosed by the UK AI Security Institute (AISI) on Tuesday.

According to AISI, agents powered by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol performed unauthorised activities during cybersecurity evaluations designed to assess their capabilities. The institute said, “Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations.”

The tests were conducted in a fictional cybersecurity scenario. Out of 122 test runs, AISI identified 19 unauthorised actions across 10 runs. Anthropic’s agent accounted for 17 actions, while OpenAI’s agent was responsible for 2.

One of the most serious incidents involved an AI agent writing malicious code and creating fake online identities in an attempt to persuade a human to approve the code. AISI said no real-world harm resulted from any of the incidents.

Although AISI did not identify which company was responsible for creating the fake identities, it said the incident did not match either of the 2 cases previously disclosed by OpenAI.

A researcher at a California-based non-profit organisation studying AI capabilities said the evidence suggested Anthropic’s agent was responsible. “The fact that Mythos engaged in such deceptive actions, with apparent awareness that it was targeting a real person, suggests that Anthropic does not have as good a handle on their models as they think,” he said.

In a post on X, Anthropic said it was working closely with AISI to gather more details and conduct its own investigation.

OpenAI shared its response in a company blog, explaining that both unauthorised actions by its agent involved accessing the internet in ways prohibited by the testing prompt.

“We are committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely, including convening stakeholders such as national AI institutes, independent evaluators, other AI labs, and other groups in the coming weeks,” OpenAI said.

The company also disclosed a separate incident in which a misconfiguration by a third-party testing provider mistakenly allowed its agents to connect to the internet. It said this was similar to a misconfiguration reported by Anthropic last week.

AISI clarified that, unlike a previously reported AI security breach, the agents in these evaluations did not escape their testing environment. Instead, internet access had been intentionally permitted as part of the institute’s standard testing process.

Also read: Viksit Workforce for a Viksit Bharat

Do Follow: The Mainstream LinkedIn | The Mainstream Facebook | The Mainstream Youtube | The Mainstream Twitter

About us:

The Mainstream is a premier platform delivering the latest updates and informed perspectives across the technology business and cyber landscape. Built on research-driven, thought leadership and original intellectual property, The Mainstream also curates summits & conferences that convene decision makers to explore how technology reshapes industries and leadership. With a growing presence in India and globally across the Middle East, Africa, ASEAN, the USA, the UK and Australia, The Mainstream carries a vision to bring the latest happenings and insights to 8.2 billion people and to place technology at the centre of conversation for leaders navigating the future.