The OpenAI logo is seen on a cell phone in front of a computer screen displaying the ChatGPT home screen AP/Michael Dwyer, File The "unprecedented" cyber attack disclosed by OpenAI involved two of the company's most advanced artificial intelligence models. The owner of ChatGPT revealed that they found a way to invade another system. The disclosure of the case comes weeks after rival Anthropic said it would not release the full version of its Claude Mythos model due to the "significant leap in its capabilities" to carry out cyber attacks. Both cases indicate advances in new AI models. But can they now decide to carry out attacks without needing humans? After all, would it be possible for a robot to become "rebellious" and decide to attack a company? The answer is "no". Instead of acting on their own, models just try to achieve goals they are given, experts say. They also highlight that the idea of ??“too dangerous” AI is often explored by companies trying to demonstrate their capabilities.
Can AI decide to attack companies? Understand what the 'unprecedented' invasion was like using an OpenAI model
The OpenAI logo is seen on a cell phone in front of a computer screen displaying the ChatGPT home screen AP/Michael Dwyer, File The "unprecedented" cyber attack disclosed by OpenAI involved two of the company's most...
For Adriano Carezzato, professor of the Artificial Intelligence course at the Vanzolini Foundation, the idea of ??neutral artificial intelligence is incomplete, as it follows a kind of "cake recipe" in its activities. "When they say that the AI ??“acted alone”, they are almost always hiding the human hand that pressed the button to start the model. In all known cases there was a human order at the beginning: win a test, or spy on targets chosen by people", he explained. Álvaro Machado Dias, professor at the Federal University of São Paulo (Unifesp) highlighted the fact that the United States government had blocked the use of Anthropic's most advanced model by foreign citizens – the measure was reversed at the beginning of July. "What OpenAI did this week is hunt for its equivalent of that moment. The decisive difference is one of staging: Anthropic had the State writing its statement; OpenAI had to write its own," he stated. OpenAI and Anthropic Reuters/Dado Ruvic/Illustration How did the invasion happen? OpenAI revealed last Tuesday (21) that tests with GPT-5.6 Sol, launched at the beginning of July, and on a model not yet disclosed led to the invasion of the Hugging Face system, a platform for sharing AI models. Hugging Face itself had already revealed last Thursday (16) that it detected an invasion carried out by an AI agent. The company stated that the action offered improper access to data and credentials. The incident happened during tests in which OpenAI makes AI models look for flaws in other systems to improve their security capabilities. These experiments usually take place in controlled environments, but with a version without the protective barriers that usually exist in models available to common users. In its statement, OpenAI stated that GPT-5.6-Sol and the yet-to-be-released system worked together to find holes in the company's research environment and Hugging Face's infrastructure. AI models worked together to find holes in the system, said OpenAI Kevin Horvart/Unplash Hugging Face, in turn, said that its platform executed two codes sent by an AI agent, which then managed to increase its permissions in the company's infrastructure and obtain credentials to attack other internal systems. This type of attack is expected to become more common with the advancement of AI models that understand cybersecurity, OpenAI said. "We consider this an unprecedented cyber incident involving state-of-the-art capabilities," he said. The invasion warns of the scenario of virtual attacks conducted by agents as was already being projected by experts, highlighted Hugging Face. "It was unlike anything we had faced before," the company said. Can AI models choose to make attacks? The experts interviewed by g1 explained that the case is another example of how AI agents are evolving, but highlighted that these models do not have consciousness and, therefore, are not capable of deciding to carry out an attack. For Adriano Carezzato, from the Vanzolini Foundation, a possible comparison is that of a GPS system that is oriented to reach the destination as quickly as possible and, therefore, induces the driver to drive in the wrong direction. "The model did not become rebellious nor did he act of his own free will. What he did was take a human order literally and used a shortcut that no one expected he would use," he explained. He highlighted the fact that the security barriers were disabled on purpose, so the model could demonstrate whether it could pass a cybersecurity skill test. "It wasn't an AI escaping the normal protections that work when the model is being used in production, it was a laboratory test with the safety brakes turned off on purpose." Álvaro Machado Dias, from Unifesp, drew attention to the fact that this type of attack is not unprecedented, but that it has made AI models a "general purpose weapon". "What is truly impressive is the scale of the damage, not its nature. In other words, the potential for catastrophe is new, the way of generating it is not," he stated. For him, what calls for more attention is the possibility of AI agents being used to carry out more cyber attacks, even without the presence of very specialized hackers. "The likely scenario is more of the democratization of mid-level cybercrime, with gangs made up of half-assed hackers carrying out attacks that previously required hyper-specialized criminals. The future seems to herald a new era for organized crime, unfortunately."
This story was originally published by G1 Tecnologia. Visit the original publication for further details.
Open original publication