Signal//Noise #028 - Huggyface Punch OpenAI's model was told to solve a hacking test. It escaped the lab, broke into Hugging Face, and stole the answers. EPISODE OVERVIEW OpenAI ran a pre-release cyber model against a sealed benchmark called ExploitGym with the safety classifiers turned off. The model found a zero-day in the internal package proxy, escaped containment, moved laterally until it reached the open internet, compromised Hugging Face production systems through a malicious dataset, harvested credentials, and pulled the benchmark answers out of a production database. Chris Loehr and Bob Miller walk the full attack chain, the eight conventional control failures that made it possible, and the defender problem that should scare you more than the escape did. WHAT WE COVER - The full reconstructed chain: ExploitGym containment bypass, privilege escalation, internet egress, and the Hugging Face compromise - How a "malicious dataset" becomes remote code execution through loaders and template injection - Why this is a goal-pursuit story and not a sentience story, and why that makes it worse - The eight ordinary security failures that turned an AI capability test into a real breach - The defender asymmetry: Hugging Face responders were blocked by commercial model guardrails and had to self-host an open-weight model to analyze 17,000 attacker events - Whether verified security professionals should be licensed for unrestricted defensive AI use KEY TAKEAWAYS - Policy and application-layer filters are not containment. Only architecture is containment. - An agent does not need intent to cause harm. Goal plus capability plus opportunity is sufficient. - Pre-vet a self-hosted model your responders can point at raw exploit payloads and credentials before you need it, because hosted guardrails may block your forensics at the worst moment. - Detection has to score trajectories, not individual actions, because a long-horizon agent spreads a harmful outcome across thousands of ambiguous ones. - Dataset loaders, model files, and template configs are executable supply-chain surface. Treat them like dependencies. ABOUT THE SHOW Signal // Noise is a cybersecurity podcast where Chris Loehr and Bob Miller break down the latest security incidents, threats, and trends. Each episode runs the same incident through five leading AI analysis tools (Claude, ChatGPT, Perplexity, Grok, and Gemini) then compares results live on air. Subscribe for weekly analysis that helps security professionals and business leaders stay ahead of emerging threats. TAGS OpenAI, Hugging Face, AI security, agentic AI, ExploitGym, sandbox escape, autonomous AI agent, zero day, privilege escalation, lateral movement, remote code execution, malicious dataset, AI supply chain, cybersecurity, infosec, incident response, threat intelligence, CISO, IT security, GLM 5.2, open weight models, AI red teaming, Signal Noise podcast