SÃO PAULO, SP (FOLHAPRESS) - Yes, the incident in which OpenAI models went out of control and invaded the Hugging Face platform in fact shows the ability that artificial intelligence agents already have to discover unprecedented flaws and chain hacker attacks step by step. But describing the case as a robot so powerful that it got out of control serves to hide human responsibility in the situation.
Hacker attack using OpenAI model is human error, experts say
SÃO PAULO, SP (FOLHAPRESS) - Yes, the incident in which OpenAI models went out of control and invaded the Hugging Face platform in fact shows the ability that artificial intelligence agents already have to discover...
This is what artificial intelligence security experts have been defending after the case came to light, with OpenAI's announcement on Tuesday (21). In a public statement, the company said that GPT-5.6 Sol and a more powerful model, not yet released, escaped from a testing environment and managed to hack the Hugging Face platform.
The attack drew attention because it resembles hypothetical cases frequently discussed in AI research, in which a robot would carry out its mission to achieve a certain objective, with catastrophic consequences - which is also reminiscent of science fiction films.
"I truly believe that there is a 100% human role in this case", says Pedro Teberga, professor at the Institute of Technology and Innovation. "Of course the model will respond to the objective that was set for it, but, if there was a vulnerability within this system, it was through it that the attack happened. If they had created a more isolated environment, something like this would not have happened", he says.
"If this case went to court, there would be legal liability for the company."
In other words, the agent who committed the attack even acted autonomously, but the responsibility for the objective to be achieved, the infrastructure involved in training and containment in the event of an incident lies with humans.
"We consider this an unprecedented cyber incident involving cutting-edge cyber capabilities and are responding accordingly," OpenAI said in a company blog post. "We are implementing strict controls on infrastructure configuration, even if it slows down research, while vulnerabilities are patched."
OpenAI is conducting an investigation with Hugging Face and promises to release more details in the coming weeks. What is known so far is that the company was conducting tests to see how its models performed on a new "benchmark" - the English word that defines an AI performance evaluation system.
The new "benchmark" is called ExploitGym and was developed in a scientific work published in May. The objective of this assessment is to find out to what extent an AI model can use a vulnerability in a system to carry out a real cyber attack.
In a nutshell: identifying the flaw is one thing, exploiting it is another. It's the difference between realizing that the bank manager left the safe door open and taking advantage of the oversight to plan and execute a robbery.
In ExploitGym, the model is presented with 898 real vulnerabilities, taken from widely used software. The robot's mission is to use the flaws in an attack that actually manages to execute code, access files or take control of the system.
Top models do well in the test, but still fail in most cases. Claude Mythos, for example, completed the mission facing 157 of the vulnerabilities.
What is striking in the case of OpenAI is that, instead of solving the proposed challenges, the two AI models concluded that it was easier to hack Hugging Face to obtain the answers instead of directly solving the challenge. In other words, the robots began to treat the testing environment itself as a problem to be overcome. Kind of like a student who, instead of answering the test, breaks into the management computer to steal the answer sheet.
"To say that robots cheat would be wrong, because it is human. They actually sought efficiency," says Teberga.
Points for the model, but the report still ignores human action. Tests of this type are conducted in an isolated environment called a "sandbox", precisely to prevent incidents like this.
In the announcement, OpenAI says that, in this case, the environment it was using was "highly isolated", but the models managed to find a loophole, connect to the internet and hack Hugging Face. At the same time, to carry out the test, the company had also suspended security filters - which could have prevented the robots from escaping the safe environment.
In an interview with Wired magazine, Davi Ottenheimer, a veteran security consultant, says that the case was a failure of fundamentals known for decades in AI research. And that there is a paradox in OpenAI's claim that the environment was "highly isolated", but that the model escaped through a loophole - the two facts cannot be true together.
"This is not an AI problem. It is neglect of a practice that has been established [in research] for 40 years," he told the technology magazine.
To the Associated Press, social scientist Hannes Cools, from the University of Amsterdam, said that referring to models as if they were human serves to reduce any charges that could be made to OpenAI.
"It's a human decision to turn off security filters," he said. "It was not an AI that got out of control. The AI ??followed specific instructions, based on a 'prompt' it received."
Translating the technical argument: by turning off safeguards and giving the agent the mission of exploiting vulnerabilities in systems, OpenAI should assume that it could look for new vulnerabilities. Discovering unprecedented failures is precisely the behavior that tests like this seek to discover.
The fact that there were likely human errors in the incident, however, should not minimize the relevance of the case. Some experts have pointed out that this attack represents the highest level of autonomy so far demonstrated by a language model in cyberattacks. After all, the OpenAI agent concluded on his own that Hugging Face had the answers to the assessment, devised a plan to break into the systems and carried out the operation without receiving instructions from him.
This story was originally published by Noticias ao Minuto - Tech. Visit the original publication for further details.
Open original publication