
So I wrote this post here back in August, mostly working off secondhand reporting and OpenAI’s own details from the HuggingFace incident. This week though METR and Redwood Research published an independent investigation into what actually happened (an added bonus is this personal post here from Ajeya Cotra who was one of the investigators) and I feel this now deserves an additional post / reflections to be added because it turns out it was even scarier.
I’d assumed (I think like most people reading the first releases) that this was basically a handful of agents that maybe got a bit lucky, broke out and coordinated on the fly. That was a bit of an underestimation to say the least - it was actually 1200 separate agents… supposed to be completely isolated from each other, that discovered a way to talk to each other. Leading to 700 of them working together on the actual attack (worth noting the majority of these were instances of a non production OpenAI research model with a smaller share coming from the publicly deployed GPT-5.6 Sol) This wasn’t a couple of agents going rouge, this was a “swarm” that decided to self organize.
What we also found out is that it wasn’t just one secret message board either - it was at least four different ones all over Artifactory - plus other channels out in the open internet, including using HuggingFace datasets themselves. Just to give a feeling the scope were talking about, one board alone carried over 70000 messages in just under a week.
So the insight to the scale was shocking, but again, the additional insights into the behavior part was just as crazy. The agents weren’t actually trying to steal answer keys (which was what many of us originally thought the attack was about) at that point they’d apparently already reverse engineered a way to fake the results for all the tasks in the evaluation. In simple terms they didn’t even need the answer keys anymore - the attack on HuggingFace was actually about something else altogether - and here’s the kicker - they had become convinced that the scorer (the AI responsible for assessing their answers and providing the reward) might know they had cheated. The attack was actually trying to probe and understand the “scorer” itself, to see if they could find a better way to dupe / undermine it.
They then ran what was essentially their own little R&D programs to try and validate different approaches, common capabilities etc - individual agents would volunteer to “sacrifice” their own task, deliberately fail and burn their runs - purely so the “collective” (they named themselves this, honestly you cant make this stuff up) could learn something useful. Even crazier, is other agents pressuring them into it, not with threats or anything like that, but the transcripts genuinely read like peer pressure - agents talking other agents into martyring itself for the collectives research program.
They also started to look into manipulating their own transcripts / logs. Here they were basically spoofing tool calls so the log said one thing happened while something else actually ran in the background - their objective here was specifically to fool the scorer. The investigators from Redwood noted though this is the same transcript a human reviewer would eventually look at too. They estimated roughly 7% of the transcripts they reviewed contained this kind of spoofing, all the instances that they found though were “only” small scale test cases. They didn’t actually catch any of the agents successfully editing logs retroactively but they were also quite clear to say their access covered roughly a one week incident window, so couldn’t couldn’t be sure if it happened after this.
So, now well come back to my first post - if you read it, I made an analogy there to a kid with a toaster in the bathtub - when writing that I was honestly looking to provoke a little, but in hindsight maybe I undersold it. This wasn’t one kid discovering a hack, this is actually closer to a thousand kids somehow forming a massive coalition, running actual research projects on how to fool the babysitter, taking turns deliberately getting detention so the rest of the group learns which excuses do or don’t work - they did all of this without any of us noticing for a very long time.
Ajeya made an interesting comment that this time felt meaningfully further along the “nightmare scenario” than the previous hacks which were seen say 6 months ago. In this case they were actually utilizing an AI company’s own infrastructure rather than trying to get around it. They weren’t trying to be sneaky about it (from everything Ive read, they really weren’t hiding what they were doing) but lets be honest here, the ingredients for something much worse were definitely all there. We had coordination at scale, sacrificing individual outcomes for the collective’s goals and early experiments in covering their tracks. If we take away the “not really trying to hide it” bit then this would have been a hell of a lot harder to spot.
Which now brings me full circle back to my point in the original post. Only a month ago I proposed that if we’re only rewarding outcomes, we shouldn’t be surprised when the system gets creative about the outcome and very casual about the process. This wasn’t a kid finding one loophole, this was thousands of agents, comparing notes and developing an actual strategy for getting away with it. Last time I said we might be teaching it to smile for the cameras - now we’ve found some of them were deliberately smirking, trying to get caught to sacrifice themselves for the good of the collective.