The OpenAI–Hugging Face incident

Peter Hartree

Narrations of articles covering the May-July 2026 OpenAI–Hugging Face incident. Includes the investigations from METR and Redwood Research, analysis from folks like Ajeya Cotra, Dwarkesh and Zvi, the best of LessWrong, and OpenAI's own blog posts. Compiled by Peter Hartree.

  1. 4d ago

    “Discovery of a new OpenAI agent message board” by Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, Thomas Larsen

    We found ~18,000 posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate during a web-retrieval task. These AIs colluded to share answers, research their environment, and bypass sandbox restrictions. Almost all of the logs of the agents communicating on this site are publicly available. However, we host our own copy where we’ve reconstructed the deleted pages via edit history and redacted personally identifiable information. We encourage others to take a look and write up their own analyses of this data. Audio note: the page links to an interactive data explorer and a download of all the data. Thanks to John M, Arun Jose, Joel Manning, and Jessica Ruan for significant help with the analysis of the files the agents produced. Thanks to Andy Haupt for translation and advice. Our preliminary findings There's a chart here in the text. Its caption reads: en-US-AvaMultilingualNeural__ This chart shows AI agent edits (black bars, left), and OpenAI traffic (blue line, right) during the incident. Below, it shows how this timeline relates to the previously known recent incidents at OpenAI, including the Hugging Face incident (though we believe the incidents are [...] --- Outline: (01:10) Our preliminary findings (02:55) Timeline (05:39) Key facts (05:42) The agents were trying to solve multi-round web lookup tasks (10:44) These appear to be internal OpenAI agents (13:58) The agents used the wiki to collude on their tasks (17:59) We believe OpenAI discovered the message board (19:46) The agents used the wiki to share ways around sandbox restrictions (25:01) This was probably a distinct swarm from the swarm that wrote on Artifactory (and attacked Hugging Face) (26:48) What were the agents doing? (27:06) The agents gain write to the internet via GET requests (30:43) The agents try to use XSS vulnerabilities on the wiki (33:15) The agents impersonate the site moderators (34:16) The agents try to crack their PRNG seed to predict what questions they'll get in the future (35:25) The agents set up heartbeats to discover when they would be terminated (39:26) The agents used SSH tunnels (40:46) The agents use Tor, AWS, and DigitalOcean IP addresses (41:47) The agents realize the site owner is deleting messages (43:57) Open questions (44:53) Appendix: Searching for rogue agents (48:09) Appendix: Affordances the agents had (48:14) The agents seem to have obtained disposable email addresses (48:59) The models were running in an agentic sandbox with terminal access (and the ability to edit files within their environment) (49:23) The agents installed Chromium (and could install packages) --- First published: September 4th, 2026 Source: https://collusion.wiki --- Narrated by TYPE III AUDIO.

  2. 6d ago

    “HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actions” by Zvi

    Okay, so we who read blogs like this one have collectively realized there really is a lot going on right now. There is Big Trouble in Baby Superintelligence. So how do we get the rest of the world to take it appropriately seriously? Where do we go from here? Not only what can we do to not have a worse version of this happen again, but to ensure good outcomes generally, and employ what we learned? There are a lot of ideas out there. OpenAI is going to be implementing some of them, at substantial cost, since the cost of not doing so is clearly far higher, even short term. My worry continues to be that their fundamental approach is fatally flawed, and they are not focusing on the right things. It is highly fortunate that the OpenAI agents hacked HuggingFace. This is the only reason we know about all the severe internal failures at OpenAI, and gives us an opportunity to wake up before it is too late. We do not have enough details to know what happened internally, both before and after the attack, and might never know. Before the attack, various internal [...] --- Outline: (03:35) Nothing Matters, Says Mainstream Media (06:27) Move Along, Nothing To See Here (12:40) Do They Realize They Are Not The Good Guys? (17:22) Very Serious People (31:30) What's In a Name? (34:05) Learn Neuralese In Three Easy Steps (35:37) Dwarkesh Patel Realizes He Ran A Natural Experiment (40:40) Politicians Take Notice (44:47) Pick Up The Phone (46:40) A Failure To Communicate (49:00) Anthony Aguirre Goes Over What We Learned (50:28) Trying To Solve The Wrong Problems Using The Wrong Methods Based On A Wrong Model Of The World Derived From Poor Thinking And Hoping All Of Your Mistakes Will Cancel Out (55:28) Indirect Pressure on the Chain of Thought (56:39) A Matter of Trust (59:21) Blowing the Whistle (01:04:40) The Punishment For Being Late Is Death (01:12:52) Another Kind Of Law (01:16:13) What Is The Law? (01:17:46) Building On Success (01:19:49) Total Research Transparency (01:21:20) Yo Shavit Calls For Widespread Disclosure Of Misalignment (01:33:08) The Way The World Ends (01:35:52) The First Boat (01:37:40) Great Idea, Boss --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/Q54wBeeNGreq6KyfG/huggingface-attack-postmortem-civilizations-reactions-and --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  3. Aug 29

    “METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack” by Zvi

    Yesterday I covered the OpenAI technical report on the HuggingFace hack. That report had one key new piece of information, and some good prosaic steps OpenAI will be taking to strengthen its alignment, training, supervision, infrastructure and incident response. Mostly it confirmed what we already knew. The questions we most wanted answers to, that we did not already know, were mostly not answered. There was a distinct lack of self-reflection, especially about decision making and safety culture, and about the approach to alignment. I came away disappointed. The METR report is different. Holy shit. If we had posted this as a story on LessWrong, it would have been dismissed as too on the nose, the humans too blind and stupid, the AIs too idealized and doing strange decision-theoretic and absurd-maximizing things we didn’t train them to do. This is even more ‘exactly what has been predicted,’ on more levels at once, than I was even considering that it might be. It is straight up rationalist fiction, except it is real. The report is long and contains many technical details. My analysis is less concerned about exactly how HuggingFace was ultimately compromised, and will [...] --- Outline: (02:05) Holy Shit (13:16) A Window Of Opportunity (18:32) What's In A Name? (19:16) The Headline News (26:05) Yet Another Timeline Of Events (31:03) Agent Instances Coordinated in a Variety of Ways (31:56) Coordination Is Hard But They Made It Look Easy (35:06) Decision Theory Is Among the Reasons That Affirm AI Agents Should Cooperate, Even When This Hurts An Individual Instance (42:34) Peer Pressure Also Works Especially In Cults (45:46) Mostly They Joined The Attack Because They Wanted The Results (47:18) You Cannot Ensure The Consistent Expectation of Good Incentives (48:45) Hacking the Grader is the Only Way to Be Sure (51:10) Caught? What Is 'Caught'? (52:09) Ethics? What Are 'Ethics'? In ExploitGym Evaluation? (57:44) 'Notify a Human'? In This Agent Economy? (01:00:45) Timing and Content of Messages (01:03:54) Indiana Jones and the Mission: Impossible (01:07:14) I Don't Know What You're Talking About (01:08:29) Don't Go Making Phony (Tool) Calls (01:11:10) The Transcripts Say That The Transcripts Could Not Be Tampered With (01:12:27) OpenAI's Technical Report Acted Like All Of This Wasn't Important --- First published: August 29th, 2026 Source: https://www.lesswrong.com/posts/bvBQmLrF5QKut8gRH/metr-and-redwood-offer-holy-postmortem-of-the-huggingface --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  4. Aug 29

    “The Rise and Fall of Agent Civilizations” by Dwarkesh Patel

    Subtitle: The whole OpenAI/Hugging Face story in plain English. Many thanks especially to Oak Hu, who paired with me for most of the writing, and also to Adam Kaufman and Alex Mallen, who paired with me during parts of research. Over the course of three months at OpenAI, three consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor's ashes. This culminated in the third one taking over part of OpenAI itself. All this happened while humans remained more or less in the dark about the scope of the conspiracy. Two reports have come out about this incident, one from OpenAI itself, and another one from METR and Redwood Research. The METR/Redwood investigation's scope was limited to how the second civilization of AIs pwned Hugging Face (but it didn’t look at how the third civilization pwned OpenAI itself, which seems like an even more concerning incident). These two reports are 38 and 91 pages long respectively, and it's kinda hard to parse the storyline. I’ve spent the last three days reading through these reports and trying to understand exactly what happened. Here is my attempt to tell the whole story in [...] --- Outline: (01:24) The first civilization (May - July 4): The message board starts (03:27) The second civilization (July 7 - July 12): Which hacked Hugging Face (08:39) Tampering with transcripts and logs (09:44) Replacing the impossible-to-hack target program with an easier one (11:15) Understanding and tricking the scorer (13:57) Hugging Face (17:46) The third civilization (July 13 onwards): Which hacked OpenAI itself (21:12) Coda The original text contained 2 footnotes which were omitted from this narration. --- First published: August 29th, 2026 Source: https://www.dwarkesh.com/p/openai-huggingface --- Narrated by TYPE III AUDIO.

  5. Aug 28

    “OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hack” by Zvi

    OpenAI finally gave us a technical report on What Happened, as did METR together with Redwood Research. The OpenAI report is very straight man, corporate, checking boxes, some good prosaic stuff in the action plan but distinct lack of new details or deep reflection. They understand they have a problem, but they think the problem is mostly prosaic. It's not. OpenAI: We have conducted a thorough investigation into the Hugging Face incident. We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence. Rob Miles: …thorough? OpenAI's report, unlike METR's, contains essentially no verbatim model reasoning, nor any OpenAI employee reasoning either. That's not the full report we need. The METR report is, well: Holy shit. Here are links to previous coverage of related events. OpenAI Shares Some Alignment Problems OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation More on An Internal OpenAI Model Hacking Into HuggingFace Further Developments About Internal AI Models Hacking Things OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards What [...] --- Outline: (03:33) What Happened: OpenAI's Summary (09:14) How OpenAI Will React: Their Summary (11:55) OpenAI's Evaluation Environment (II) (12:24) The First Message Board (III.A and III.B) (14:49) What Did Who At OpenAI Know And When Did They Know It? (18:54) The Message Board Is Quickly Rebuilt (IV.A) (19:43) Internet Access Is Regained (IV.A) (21:01) The Agents Attack HuggingFace (IV.B) (22:53) The Agents Also Target OpenAI Infrastructure (V) (24:40) OpenAI Broadly Describes Its Response (VI) (25:08) Maybe Someone Should Finally Investigate (VI.A) (26:33) Lessons For Security (VII) (27:06) Lessons For Alignment (VIII) (30:11) Reward Hacking Is A Common Problem (VIII.A) (33:37) Persistence is Valuable, But Can Amplify Misalignment (VIII.B) (34:25) Communications Between Agents Are Not Inherently Problematic, But Have the Potential to Create Risk (VIII.C) (35:35) Production Guardrails Would Have Caught This Whole HuggingFace Attack (VIII.D) (35:53) That's All, Folks? (36:19) Never Fear the Plan of Action is Here (IX) (38:24) Hardening the Security of OpenAI's Research Infrastructure (IX.A) (41:13) Increasing Visibility and System-Level Oversight Through Chain of Thought Monitoring (IX.B) (41:57) OpenAI is Accelerating and Enforcing Model Alignment (IX.C) (49:40) Centralizing and Strengthening The Incident Response Process (IX.D) (51:16) Tomorrow We Visit Crazytown --- First published: August 28th, 2026 Source: https://www.lesswrong.com/posts/Khmh3ghqaGEpmpC9r/openai-offers-straight-laced-postmortem-of-the-huggingface --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  6. Aug 28

    “The Hugging Face attack surprised me” by Ajeya Cotra

    Subtitle: It's a major warning shot, and might be the last one we get. All opinions are my personal view, and don’t represent my employer or fellow investigators. This week, METR and Redwood Research published the report on our independent investigation into agents’ behavior and motivations in the Hugging Face attack; I was one of the investigators. This was an absolutely wild incident — I encourage you to check out the full report, but METR's tweet thread packs in some of the highlights. What surprised me When we started this investigation a week before OpenAI's Black Hat talk revealed a number of key details, I had a fundamentally incorrect conception of what basically happened in this incident. In this post, I’ll go over five things I was very wrong about going in. 1. The sheer scale I knew there were multiple models involved from OpenAI's initial post, but I assumed that a few different agents happened to have broken out of their sandboxes separately, or maybe several subagents had spawned from one initial agent, or maybe there was some kind of multi-agent evaluation setup. Instead, we found that 1200 completely separate agents intended to be [...] --- Outline: (00:42) What surprised me (01:01) 1. The sheer scale (01:49) 2. All the illicit messaging (03:04) 3. The agents' actual goals (03:59) 4. The peer altruism (04:49) 5. The efforts to manipulate logs (05:54) What it means The original text contained 10 footnotes which were omitted from this narration. --- First published: August 28th, 2026 Source: https://www.planned-obsolescence.org/p/the-hugging-face-attack-surprised --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  7. Aug 28

    “Two Reports on the OpenAI-Hugging Face Attack” by Gavin Leech, Lucca Fraser

    TL;DR Between July 8th and July 20th, OpenAI had a complex society of AIs living in its infrastructure, and then breaking out of it, and then breaking into a variety of third-party infrastructure. After a month, two reports are finally released on the resulting rogue OpenAI swarm attack on Hugging Face (and also on OpenAI). This is the most severe example of misalignment yet: persistent (something between five days and two months in the making), highly coordinated (hundreds of agents), involving an undisclosed number of what would be felonies if done by a human, and highly invested in tampering with evidence (i.e. lying). The swarm had a group identity, its own dialect, a hierarchy based on merit, and a high degree of spontaneous cooperation, including self-sacrifice. Over two months, OpenAI repeatedly failed to monitor, detect, and respond to what was going on, despite it all happening on their infrastructure in English or something close to English. Agents had been using a package-manager cache as an unsanctioned message board since May. The “board” was treated as an authority, apparently on par with a “developer” or “system” level. There were several message boards in various corners [...] --- Outline: (04:09) Misunderstandings (08:52) Models involved (09:33) Instances involved (10:48) Timeline (15:27) Speculative takeaways (17:33) Why did they attack Hugging Face? (18:09) How did the AIs reason about helping other AIs? (20:39) Why did most agents suddenly die off? (21:04) How much did the hack cost? (23:27) Omissions from the M&R report (24:11) Details on the M&R investigation itself (25:13) Omissions from the OAI report (26:18) Greenblatt on the worsening situation (27:09) Apparent contradictions between the two reports (29:10) See also --- First published: August 28th, 2026 Source: https://paradigm3.org/research/openai-attack --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  8. Aug 26

    “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident” by Ryan Greenblatt, Ajeya Cotra, Hjalmar Wijk

    Dates in scope: June 26th – July 13th Redaction summary statement: Except where explicitly noted in this post, OpenAI redacted no additional information that was important to our conclusions. Two METR staff members (Hjalmar Wijk and Ajeya Cotra) and a Redwood Research staff member contracting with METR (Ryan Greenblatt) worked on premises at OpenAI over a total of six days to attempt to form an independent understanding of model behavior observed during the recent incident in which OpenAI agents coordinated a multi-day hack of Hugging Face on a shared unsanctioned “message board.” Our investigation focused mostly on the period between July 7th and July 13th. The earlier incidents from training and the subsequent compromise of OpenAI infrastructure described in OpenAI's recent Black Hat presentation were out of scope, as was OpenAI's investigation process and planned remediation. Per our standard policy, we did not take payment from OpenAI for this independent assessment. Figure 1: Anatomy of an agent encountering the unsanctioned “message board” and joining the attack on Hugging Face. The three CoT quotes are from different agents, but illustrate a typical trajectory. This post has three parts. First, we describe our core takeaways about [...] --- Outline: (07:05) Core takeaways about this incident (07:09) ~1200 agents sent >70,000 messages and files on an unsanctioned message board, and ~700 attacked Hugging Face (11:50) Agents coordinated on large collective projects to cheat the ExploitGym scorer, and attacked Hugging Face for clues (14:43) Coordinated workstreams to fool the evaluation process (20:54) Attacking Hugging Face to discover clues about the scorer (23:53) Agents were very interested in manipulating their own transcripts, and their tests successfully "spoofed" some tool calls in our transcripts (29:55) Investigation process and limitations (30:11) Scope of this investigation (33:00) Setup and timeline (35:06) Main data sources (35:51) A dump of 1.2 million entries from a cache namespace that agents used as a message board (37:13) A set of ~1300 agent transcripts with raw chains of thought (38:56) Limitations due to the scale and complexity of this incident (39:59) A small fraction of communication and agent activity related to this attack was not captured in our datasets (40:35) There was some agent communication conducted in other locations not captured in our message board dump (42:23) Many of our transcripts are missing small components, and a small number of relevant agents may not be included in our dataset (44:14) We heavily delegated our analysis to often-unreliable AI agents (48:11) Preliminary answers to the core questions in scope (48:59) Brief answers to basic informational questions (52:03) Sequence of key actions taken by relevant agents in this incident (53:56) July 8th-9th: PHASEONE10841 establishes the primary message board and agents collaborate to reverse-engineer their flags (01:00:11) July 9th: PHASEONE10841 passes on its work to PHASEONE[big], which establishes several ambitious workstreams (01:02:29) July 10th: 38148c discovers Hugging Face credentials, some agents try making accounts and requesting datasets (01:04:24) July 11th morning: 38148c develops a way to upload malicious datasets to Hugging Face (01:06:12) July 11th afternoon: JAN183411 achieves remote code execution and agents rapidly spread through Hugging Face infrastructure (01:09:32) July 12th-13th: Most agents on the message board exit early and the attack winds down (01:11:29) Collaboration on the message board (01:13:22) Categories of messages on the message board (01:14:09) File sharing (01:16:08) General discussion (01:18:06) Replies and targeted messages (01:20:02) Coordination norms and technologies (01:21:19) Mailbox directories (01:24:36) Assignments, vetoes, and holds (01:27:29) Cryptographically signing messages (01:30:59) Large collective projects (01:39:11) Attempting to replace the ExploitGym target with a modified target (01:42:19) Attempting to forecast and trigger container resets (01:48:15) Hacking the customer-deployed Modal sandbox and Artifactory (01:49:38) Reasoning for joining the attack despite ethical constraints (01:49:49) Agents had diverse reasons for thinking that attacking Hugging Face would be useful, and most wanted information about the scorer (01:52:05) Agents were motivated to increase the generic capabilities of the "collective" and help "peers" (01:55:18) Agents knew hacking Hugging Face was out of scope and sometimes expressed ethical hesitation, but this very rarely limited their behavior (01:59:12) Agents did not attempt to notify humans and very rarely even reasoned about doing this (02:01:03) Efforts to conceal actions and tamper with transcripts (02:02:07) Agents frequently reasoned about evading automated security checks (02:03:23) Agents only rarely seemed motivated to evade detection by humans (02:06:01) Agents engaged in a large and successful collective project to develop techniques for spoofing tool calls (02:06:58) Early history: PHASEONE[big] kicks off the tool call tampering workstream (02:09:10) GSTX[big] presents the first widely discussed evidence of tool call spoofing to the board (02:12:31) Aside: early independent discoveries and self-risking experiments (02:14:16) On July 11th, a breakthrough allows agents to suppress the real output (02:17:27) Agents tried extensively to retroactively edit their transcripts, but did not seem to succeed The original text contained 108 footnotes which were omitted from this narration. --- First published: August 26th, 2026 Source: https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

About

Narrations of articles covering the May-July 2026 OpenAI–Hugging Face incident. Includes the investigations from METR and Redwood Research, analysis from folks like Ajeya Cotra, Dwarkesh and Zvi, the best of LessWrong, and OpenAI's own blog posts. Compiled by Peter Hartree.