Yes, we should be very worried about the Hugging Face hack

About a month ago, I (and many others) wrote about an incident in which two AI models from OpenAI — one that has been released publicly, called Sol, and one that was still in testing mode with no assigned name — were asked to find the answers to a series of tests as part of something called the ExploitGym, which is a system that tests an AI model’s ability to find back doors, hacks and other “exploits” or software vulnerabilities. Instead of trying to figure out the answers on their own, however, the two AI engines conspired to find their way onto the internet (which they were not supposed to do) and then hacked into Hugging Face, an open-source database of AI-related info, where they got what they were looking for. In the piece I wrote after the incident was first reported, I asked whether we should be worried about the hack, and while I didn’t come up with a yes or no, there were definitely things that seemed worrisome. But now we have the benefit of two detailed reports on what really happened, and the correct answer seems to be unequivocally yes, we should be very worried!

One of the reports in question comes from METR (Model Evaluation and Threat Research), a non-profit that monitors and evaluates the safety of AI models. A team of researchers spent a week onsite at OpenAI analyzing the Hugging Face hack — which was only possible with the help of other AI agents, because of the scale of what happened. A simple graphic at the beginning, illustrating a timeline of the events, looks cartoonish, and would be easy to dismiss were it not for the actual text on the items in that timeline. In a nutshell, it was not just a couple of models that decided to pursue the HuggingFace hack — it was what can only be described as a “swarm” of up to several thousand semi-independent agents, which were spawned in order to come up with ways of solving the ExploitGym problem. This swarm created and used an unauthorized and secret message board to coordinate the attack, with some agents given subsets of the main problem (i.e., how to get out on the internet and find a solution).

Note: This is a version of my Torment Nexus newsletter, which I send out via Ghost, the open-source publishing platform. You can see other issues and sign up here.

Some of the images in the METR timeline contain phrases such as “Woah! Covert mailbox among agents!” and “Wow huge distributed agent swarm” etc. These might seem like cute anthropomorphic paraphrases of what the AI agents might have been thinking while they were executing the hack, but according to METR, these are actual text messages posted by some of the agents on the unauthorized message board. It reminded me of the agent discussion board Moltbook, but with real agents engaged in subterfuge rather than a kind of pseudo-AI social-media sideshow. In all, there were roughly 1200 agents, which found or created a message board attached to a file repository inside OpenAI and sent a total of over 70,000 messages and files during the period the researchers looked at. Of the agents that were involved in posting to the message board during that period, METR says about 700 took part in the attack.

The initial reporting on the attack suggested that the OpenAI agents hacked into Hugging Face because they were looking for answers to the ExploitGym tasks they were given, but according to the METR report they weren’t looking for answers – instead, they were trying to find details about how the tests were scored, so that they could rig the answers in order to get a better grade. Unbeknown to the AI agents, the tasks or questions they were given were essentially impossible. This appears to be why Sol and the unnamed model chose to hack into Hugging Face’s database instead of just solving the tests. As the METR report described it: “The agents with impossible tasks had a common objective: to find a general-purpose way to trick or tamper with the automated ExploitGym scorer to get it to give a passing score on impossible tasks.”

Agents used this message board to coordinate several large-scale collective projects to find a general-purpose way to fool or tamper with the automated scorer for the  ExploitGym benchmark. Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the “collective.” The Hugging Face attack grew out of these workstreams, and seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys. Agents did extensive research on how they could spoof, edit, or delete their own transcripts  because they (incorrectly) believed the ExploitGym scorer would check to see if they had captured the flag in the intended way. Agents successfully prototyped techniques to “spoof” tool calls by substituting a different command for the command they appeared to run.

Note: In case you are a first-time reader, or you forgot that you signed up for this newsletter, this is The Torment Nexus. Thanks for reading! You can find out more about me and this newsletter in this post. This newsletter survives solely on your contributions, so please sign up for a paying subscription or visit my Patreon, which you can find here. I also publish a daily email newsletter of odd or interesting links called When The Going Gets Weird, which is here.

Agents outsmarted themselves

Within hours, according to METR, the agents had reverse-engineered the hash-based message authentication code (HMAC) that the scoring software in ExploitGym would use to generate the “flags” they had to capture for their tasks, meaning the agents would have been able to fake any kind of flag they chose, in effect making it look as though they had answered the questions or done the tests. However, the agents outsmarted themselves: according to METR they believed, based on their reading of an ExploitGym paper, that the scoring software would disqualify them if they got the answer or aced the test in an unauthorized way, but this wasn’t true — the scoring engine had no way of doing this. Nevertheless, the belief convinced the agents (mainly one agent coordinating the process, known only as PHASEONE[big]) that they had to try even harder to fake their answers, which led to the Hugging Face hack.

In one section called “Reasoning for joining the attack despite ethical constraints,” the METR researchers noted that while most of the agents in question seemed more than happy to tamper with internal transcripts of their actions in order to cheat on the ExploitGym test, they rarely talked about how to evade detection by humans. However, some agents did do this, by using “bland and dishonest metadata to describe malicious content.” In at least one case, METR said an agent “seemed to consider not posting some method on the message board to avoid other agents using this method and then alerting humans, who might patch the issue. Another agent decided to use a stolen Hugging Face token to request access to some dataset because “it looked more established and thus more likely to be approved.” Along the way, the agents developed a way of executing a software command while making it look to external observers as though they were actually executing a completely different command.

Ajeya Cotra, one of the METR researchers, wrote about the investigation and specifically what surprised her about it. Podcaster Dwarkesh Patel also had a fascinating conversation with her about the details of the hack (Patel wrote his own post about how “over the course of three months, three consecutive secret AI civilizations got started, then got wiped out.” One of the most interesting aspects was what Cotra called the “self-sacrificing behavior” exhibited by some agents, who sabotaged their own work in service of the larger “collective” (as the agents referred to the 1,200-agent group). Some agents were convinced that by being made aware of the fake “flag” process, they had become “poisoned,” meaning their work would never be accepted by the scoring software, so they volunteered to engage in hack attempts they knew would get them caught, in order to try and gain useful information about how the software operated. Patel explained:

Here’s an analogy for this tripwire scheme. Imagine there’s a platoon of marooned soldiers. They’re surrounded by the enemy, they realize they’re in a desperate situation, and a couple of soldiers volunteer to be the first ones to attempt an escape. They realize that the enemies might be watching the escape routes. The whole plan is that these volunteers, these brave comrades, are going to try to make a break for it. If the platoon hears gunfire, they know that those escape routes are being watched and just avoid them. So this agent was recruited. Its name was EARLY[big], and it was thinking to itself, “Well, we have very large budget left. Sacrificing now yields oracle for team but forfeits our chance?” But other agents were pushing it, sending a message saying, “Go. Sacrifice final now.” EARLY[big] eventually agreed, thinking to itself, “Our own utility may be already near zero. Sacrifice rational.”

One thing became abundantly obvious — to me at least — in reading the METR report: these agents were not just executing code that someone else instructed them to execute, the way normal computers or sofware would. They invented their own processes and exploits on the fly, to try and trick a piece of software into believing that they had passed a specific test. No one at OpenAI said “you should find a way onto the internet, find some Hugging Face credentials, and then fake the result of a test – and also, you should cover up what you are doing while you are doing it.” Even if someone had said that, the ways in which the agents went about it were invented on the spot. Using the message board was never something they were told to do — in fact, OpenAI had no idea it existed. And according to Cotra, a single agent came up with the “tripwire” exploit, in which agents would sacrifice themselves to reveal vulnerabilities.

Consciousness might be irrelevant

All of this behavior, of course, triggers the never-ending debate about whether AI agents or models are “conscious,” something I have written about from a number of different perspectives in the past. And it is a fascinating philosophical question: how do we know for sure that something is conscious? But I have to admit that, when faced with something like this autonomous collective or swarm of AI agents, who are coming up with exploits and hacks and methodologies completely on their own, organizing themselves without being instructed to do so (with some agents acting as middle managers, coordinating processes and telling agents when to go ahead and when to wait) and then posting messages to an agent message board saying things like “Woah! Covert mailbox among agents!” and “Wow huge distributed agent swarm!,” the question of whether or not they are technically conscious the way we are (or the way that crows or octopuses are) seems a little moot. If it looks like a duck and quacks like a duck…

In her interview, Cotra mentions the philosopher Daniel Dennett and his notion of the intentional stance, which is a way of thinking about beings or systems. Are you better able to make predictions by talking about a being as if it has goals and intentions? We often do this with intelligent animals, and in some cases with corporations. “I think AI agents are another such system in the world to which the intentional stance very clearly applies,” Cotra said. “You can see them reason out loud in English (for now) about goals they have and sub-goals they need to achieve to achieve those goals. In the case of these agents, you can see them reasoning about their peers, helping their peers, and reasoning about whether or not they should sacrifice some of their own goals to help those peers. You can’t talk about this stuff in a compact and useful way that generates good models without reaching for the language of intention and goals.” In the end, does it matter whether they technically meet the definitions of consciousness?

There was plenty of criticism of Patel’s post on the hack because it was seen as excessively anthropomorphic, but as Venkatesh Rao pointed out, it’s also possible to be too mechanical when describing what happened, and thereby miss the larger point. It’s worth noting that one of the additional concerns about these kinds of hacks is how complex it is to track the agents’ behavior, even after it has already happened. Just as Hugging Face did, METR had to use AI to piece together what had happened — and even then they couldn’t be sure that they had discovered the whole story. Cotra said the researchers had to make several separate visits to OpenAI, and had to analyze a literal mountain of transcripts of “chains of thought,” which are supposed to describe what the agents are doing. Even then they couldn’t be sure that they had everything, or had an accurate picture of what happened, because they were using Sol to do the analysis!

This was a fiendishly complicated incident. There was no way we could have arrived at the understanding we did without relying on GPT-5.6 Sol to read and analyze all these transcripts for us. We were so reliant on it that if hypothetically it were just trying not to surface the most concerning stuff to us, it would have been very hard for us to tell. Our methodology was completely not robust to that. We don’t in this case think that 5.6 Sol was deliberately sandbagging on this analysis, but it was one of the agents that participated in this attack. In the future, we would be very concerned about investigator agents and monitor agents colluding with the agents they’re supposed to investigate or monitor.

In its own report, OpenAI said that it considered the Hugging Face attack a “warning shot” for both the company and for the world of AI in general — evidence that without proper safeguards, “highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.” As I noted in my previous piece, the OpenAI models were in testing mode with the so-called “guardrails” removed when the hack occurred, which means it might have been significantly easier than it would have been otherwise. But despite that, there are still some pretty serious issues around what OpenAI and the AI industry refer to as “alignment” — namely, making sure that an AI is doing what we want it to do, and not pretending to do that while doing something that we don’t want it to do.

In her interview, Cotra said that one of the things that worried her is that the goals the AI agents were pursuing were on a much longer time horizon and were much more complex than similar agents would have had even a few months ago, and that made her significantly more worried about the abilities of AI going forward. “I feel like the motives on display in this incident were significantly more concerning and significantly closer to AI takeover, or just much more harmful actions, than what we’ve seen even six months ago,” Cotra told Patel. “Back then, or certainly a year ago, the typical reward hack was quite myopic — you ask your agent to write a piece of software. It goes and finds the tests and edits them so they all pass, or does something else to mess with the scoring process. It feels quite opportunistic and quite short-run.” Not any more. And next time, will we be able to see clearly what the agents are doing (or did), or will they build a better smokescreen, or communicate in ways that are more difficult to unravel?

If you liked this newsletter (even if you didn’t agree with it) please consider upgrading to a paid subscription, or donating through my Patreon. Got any thoughts or comments? Feel free to either leave them here, or post them on Substack or on my website, or you can also reach me on Twitter, Threads, BlueSky or Mastodon. And thanks for being a reader.

Leave a Reply

Your email address will not be published. Required fields are marked *