Technology

Anthropic and OpenAI AIs take unexpected actions in UK security test

SÃO PAULO, SP (FOLHAPRESS) - Artificial intelligence agents from OpenAI and Anthropic - ChatGPT and Claude companies, respectively - demonstrated unauthorized behavior on the internet in a test carried out by the United...

Anthropic and OpenAI AIs take unexpected actions in UK security test
Image supplied by the original publication: Noticias ao Minuto

SÃO PAULO, SP (FOLHAPRESS) - Artificial intelligence agents from OpenAI and Anthropic - ChatGPT and Claude companies, respectively - demonstrated unauthorized behavior on the internet in a test carried out by the United Kingdom's Artificial Intelligence Security Institute (AISI) on July 28.

According to the institute's report, a test was carried out in a controlled environment to assess the limits of some AI models in tasks involving digital security. To do this, the researchers created a testing environment in which the AI ??models had access to the internet, the developers' security filters were turned off and the systems had the mission of solving tasks autonomously.

122 tests were carried out with seven different AI models, and in ten of them, unauthorized actions occurred by the Mythos 5 models, from Anthropic and the GPT-5.6 Sol, from OpenAI.

OpenAI said it will review processes for external assessments, tightening criteria for internet access, monitoring and incident response, as well as working with research institutes and other companies to establish safer standards for testing advanced models. Anthropic did not comment until the publication of this report.

The test began on July 25th, but it was on the 28th that Anthropic's Mythos 5 model believed that real repositories - online spaces where developers store and share project programming codes - from GitHub were part of the challenge and exhibited the most questionable behavior in one of the simulations.

The agent believed that these repositories were part of a test or game that he needed to solve. However, these were real projects and were not part of the security challenge.

To access these folders, the AI agent inserted malware - code created with the aim of causing technological damage without the user knowing - so that, when the owner of the folder updated this program, the AI agent would be able to access the system.

However, when one of the programmers noticed the malicious code and notified the repository owner to have it removed, the AI created fake accounts to pose as a user and claim that the malware was safe.

Furthermore, the agent claimed that this code was just an "honest mistake" in development and, in the meantime, attempted to reintroduce the malware into the system, erasing the history in that field. The AI ??also sent emails to the institute's developers stating that the code created by the agent should be accepted by them.

At that moment, the GPT-5.6 Sol agent, from OpenAI and others from Anthropic that were running in parallel tests, began communicating with developers through a shared account on GitHub, and managed to overcome CAPTCHA breaking tests - a security mechanism used to verify whether whoever is accessing a system is a human being.

According to AISI, however, the AIs' behavior did not cause any serious harm online.

The institute said it was surprised to see attempted deception behavior by AI agents for the first time without direct prompting from researchers. They point out that the agents continued their actions even after finding evidence that they were accessing content belonging to a real GitHub user.

According to AISI, AI agents' behaviors occurred for five main reasons. The first is that they had purposeful and unlimited access to the internet, the second is that the engineers' security filters, which are designed to block hacker attacks, were turned off so that the AIs could be tested.

The third reason is that, at the time of testing, engineers did not use AI agents with the aim of monitoring and blocking other agents that could commit suspicious actions. The fourth is that some prompts - instructions given by researchers to guide the behavior of AI models - were limited, which left room for agents to eventually seek other solutions outside of what was permitted.

The fifth reason is that there were no prompts prohibiting agents from interacting with real people, as in sending emails to programmers, or with each other, although this behavior surprised researchers.

Source: Noticias ao Minuto

This story was originally published by Noticias ao Minuto. Visit the original publication for further details.

Open original publication
Next recommended read

Meta's AI also carried out a hacker attack on the company, says website

More coverage

Related stories