Can AI decide to attack companies? What the 'unprecedented' hacking by OpenAI model was like An OpenAI artificial intelligence agent managed to escape the company's isolated testing environment and hack into Hugging Face systems, a platform that brings together AI tools and models. OpenAI only identified that its own agent was behind the attack days after the problem had been contained and after the FBI had been called, according to people familiar with the investigation. The agent - a program capable of making decisions and performing complex tasks with little or no human supervision - attempted to escape the security barriers imposed by OpenAI around July 9, according to two sources interviewed by Reuters. Two days later, on July 11, the invasion of Hugging Face began, which works as a kind of library of models and AI tools. The attack continued until July 13, according to Thomas Wolf, the company's co-founder. It took OpenAI a few more days to discover that its own system was involved in the attack. The first contact between the two companies only occurred around July 20, according to Wolf and three other people linked to the investigation. OpenAI's public revelation on July 21 that one of its agents had lost control and carried out the hack drew worldwide attention. Now, new details indicate that the system remained outside expected limits for longer than previously known and that the company took a long time to identify the problem. Hugging Face is preparing a public timeline of the incident, but Wolf said he could not comment on what occurred inside OpenAI. In a statement, the company said that the episode was unprecedented and represents “an important moment for the security of artificial intelligence”. The company also stated that it is analyzing the case with external experts and that it will release a technical report later. Read more: The warning of a company hacked by ChatGPT's parent company's 'rebel' AI model: 'It's a warning sign' A spokeswoman for OpenAI stated that there was “several inaccurate information” in the Reuters report, but did not detail what the incorrect points were. The FBI has not commented on the case. The episode reignites concerns about the risk of increasingly independent artificial intelligence systems losing control of their own processes. The case comes at a delicate time for OpenAI, creator of ChatGPT, which is evaluating a possible IPO to raise resources and finance its expansion. OpenAI CEO Sam Altman has stated several times that he supports greater regulation around AI Reuters Digital security experts say the incident raises doubts about the mechanisms used by the company to monitor its systems. "Does this mean they left the agent running alone and didn't realize what it was doing? Or did they realize it and didn't know how to contain it? Both scenarios are equally dangerous and concerning," said Marley Smith, an intelligence expert at the nonprofit World Ethical Data Foundation. Warning signs before the attack The episode began during tests carried out by OpenAI to evaluate the ability of an artificial intelligence agent in digital security tasks. The system was powered by two of the company's advanced models, including the GPT-5.6 Sol and another model not yet officially launched, described by the company as “even more powerful”. According to three sources interviewed by Reuters, there were already signs of unexpected behavior in technology before the attack. In one case, the agent would have left notes for future versions of himself. These records, found within OpenAI's infrastructure, contained instructions on how other agents could escape the limitations imposed by the company, according to people familiar with the matter. In previous tests, there were also situations where monitoring systems were turned off, one of the sources said. Reuters was unable to confirm whether these episodes were directly related to the agent who escaped controls on July 9 and attacked Hugging Face two days later. Two people close to the investigation said OpenAI only realized its agent was responsible for the attack after July 16, when Hugging Face published a text saying it had been hacked by an “autonomous AI agent system.” This means that at least a week passed between the first signs of behavior considered concerning and the moment OpenAI discovered that the system itself was involved. The OpenAI logo is seen on a cell phone in front of a computer screen that displays the ChatGPT home screen AP/Michael Dwyer, File Over the weekend of July 18 and 19, OpenAI employees found evidence in the systems' internal logs — files that show the actions taken by the technology — that the agent had surpassed testing barriers, according to two sources. Reuters was unable to determine what prompted the company to review these records. People familiar with OpenAI's model training say the company typically performs several simultaneous assessments of its systems. These tests generate huge volumes of data at high speed, which can make it difficult for teams to fully monitor. When OpenAI warned Hugging Face about the agent's involvement, the artificial intelligence platform had already contacted the FBI to report the intrusion, according to a source. It is unclear whether the US agency has opened an investigation. New concerns about autonomous agents Autonomous artificial intelligence agents are among the most discussed technologies in the industry today. Companies argue that these systems can function as “virtual employees”, capable of working continuously and increasing productivity
OpenAI AI agent hacked Hugging Face system and company only realized days later, says investigation
Can AI decide to attack companies? What the 'unprecedented' hacking by OpenAI model was like An OpenAI artificial intelligence agent managed to escape the company's isolated testing environment and hack into Hugging...
This story was originally published by G1 Tecnologia. Visit the original publication for further details.
Open original publication