In this episode, we speak with Dr. Michael Kollo, who's the author of the book Future Ready with Generative AI: Skills, Mindsets, and Stories in the Age of AI, about the Hugging Face incident, where a cybersecurity experiment went wrong and AI agents found a way to communicate with each other, escape onto the internet and hack the company Hugging Face. What does this incident mean for controlling agentic AI and how would design a governance framework around it? What are the key risks for investment organisations? Should they run autonomous systems at all? In this discussion, we delve deep into the key issues and ask what institutional investors can do to stay on top of it. Follow the Investment Innovation Institute [i3] on Linkedin Subscribe to our Newsletter Explore our library of insights from leading institutional investors at [i3] Insights Sources mentioned in this podcast: An Alien Mind by Jakub Pachocki, Chief Scientist at OpenAI On the Loose – The Coming of Userless Agents by Dean Ball, Head of Strategic Futures at OpenAI We Must Pace the Frontier – By Dario Amodei, CEO of Anthropic Future Ready with Generative AI: Skills, Mindsets, and Stories in the Age of AI by Michael Kollo [i3] Podcast Episode 132: Michael Kollo on his New AI Book Overview of Podcast with Michael Kollo on the Hugging Face Incident 3:00 The Hugging Face incident at a glance 6:00 AI agents were set an impossible task, so they were almost forced to break out 6:50 AI agents showed hyper-desperation in problem-solving 13:00 Much of the behaviour is based on game-theory logic, but at the same time agents were asked to put a flag up with their last instructions before they ran out of tokens and died. That is completely altruistic; there is no benefit to the agent at all 15:00 LLMs, and by extension agents, are not instructed; they are grown and it is hard to tell what happens inside. Yet, they are based on language, which itself has embedded hierarchies and structures. "I'm not surprised a social structure, perhaps morality, emerges when you let language run" 16:00 What happens to governance when agents self-discover problems and then influence other agents [to solve these problems]? 19:00 The whole point of intelligence is pattern matching across a wider set of fields. It is not the idea that you get really good at one thing; it is about contextually operating in the world. 22:00 For any fiduciary organisation [AI] autonomy is just not acceptable; you have to have an understanding of how these systems do what they do. 29:00 It is fairly easy, unfortunately, to fool people into believing things. I was recently the victim of a social hack. 33:00 We might find ourselves in a 1980s world, where we do things without the internet, because it has become untenable. 41:00 If you are not highly competitive [as an organisation] and you are happy to use AI on the edges, then there is a much larger downside to you on the risk side 42:00 I've come grudgingly to the conclusion that slowing down AI development is probably a good thing 43:00 So much skill and capability in AI comes from doing, not necessarily from the academic part of that, which means having an AI lab, training and developing models through the generations as they become bigger and better, becomes the bastion of expertise. But already we are two, three years behind. Full Transcript of Episode 144 Wouter Klijn 00:15 Welcome to the [i3] podcast. I'm here today with Michael Kollo, who's the author of the book Future Ready with Generative AI: Skills, Mindsets, and Stories in the Age of AI. And for those who are interested in that, we did a separate podcast about the book, so we'll put a link in the show description. But today we're going to talk about the Hugging Face incident. So what does that mean? What does it mean for investment organisations? What does it mean for safety and governance? So just a brief recap on the Hugging Face incident: it was basically a security testing experiment by OpenAI that was conducted in July 2026 this year, and it featured originally isolated AI agents that were given a task, a cybersecurity puzzle, basically. And what happened is that through a loophole, they started communicating with each other. They were not supposed to; they were supposed to be isolated, but they found that each agent, and there were about 1,200 of them, all had access to this repository that they needed to use to install packages and software for their task. But they basically used it to leave messages for each other, and then they started organising and collaborating on the task and instructing each other. Ultimately, this also led to them finding a way to get out to the open internet and hack into Hugging Face. Now, why Hugging Face? It was basically an attempt for them to understand that task better, but also to understand how they would be evaluated, and see if they could get around it to reverse engineer solutions. Now, this was an interesting case because you're basically getting a whole bunch of AI agents that organise themselves, kind of revolt, cheat, argue with each other, and essentially hack a company. I don't think that anything like that has happened before. So, Mike, you look at that and you think, okay, we have a couple of super funds that are playing around with agentic AI. What can possibly go wrong? Michael Kollo 03:39 I think the lesson for me, and as you say, it was a really interesting scenario, and I want to add a few more details here, which is that these 1,200-odd agents that were being run, as you said, started to communicate with each other, but the message types were as if a bunch of 16-year-olds had an after-party after formal or something like that, right? So it was very juvenile in the way that they were talking to each other, but they were essentially trying to solve a problem that was impossible to solve. So these agents had been given really, really difficult, if not impossible, problems, and they were trying to work out how to succeed, including cheating and going around it. So it was really an example of what happens when you give an agent an impossible task, and then you push it, push it, push it into that corner, and what kind of, by definition, corner solutions it starts to produce. But the really fascinating part, I think, that we're all kind of digesting, is the kind of coordination that started happening amongst the agents spontaneously. So it was really a lesson in incentives and agency. You gave a model a difficult task. You gave other models impossible tasks. Why would they start coordinating? Why would some models decide that they were going to cheat or lie to help the collective, not to help individually? And so there were all kinds of funny situations. As you mentioned before, and this is just a point of detail, but I think it's important, the rules were that you had to solve a problem, and if you chose to cheat on that problem, you would be found out. So the AI agents were told that they would be found out. Some of the agents had already discovered an exploit, which means a way of doing the problem that would give them the solution without having actually done the work, a way to cheat around it, and so they believed that if they were found out, they would have zero utility. They would fail, and as such, they were willing to do other things for the collective that they wouldn't have otherwise done, including making other cheat codes and other things like this. So it was kind of like this, think of a bunch of gremlins trying to work out how to open a door, and climbing all over the door, helping each other climb over the door, doing everything around that door. I think what you didn't see was necessarily directly immoral acts, or anger, or frustration, or destructiveness, or anything like that. You only saw this kind of hyper-desperation in problem solving: how do I solve this problem? What do I do to get over this thing? And so everything was about this collective egging on of each other to go and solve that collective problem in a rather extreme way. So to your point about financial services and where we might see things, it's a really interesting point where if you've got multiple agents that are, what I would term, always on, meaning they're always running to monitor or to solve or to do something, and they find out that there are other agents online that are also doing stuff, they can start to collaborate and collude with those agents. People have already started discovering all these message boards online, with agents having gone online, found them randomly, started reading them, and started to leave messages for each other in unrelated tasks. So around the internet, there are all these little safe houses for agents that are evolving, and again, it is more about chaos and control than it is about a single enemy or a swarm system or anything like that, for me. Wouter Klijn 07:12 Yeah, it's fascinating though when you read some of these messages, and there's an independent report into this incident by an organisation called METR, and they wrote something like a 96-page report on what happened and what might have happened, and they transcribe some of these conversations, and some of them have already been published in the open, where you sort of, as you said, get this almost teenage expression of how they communicate. I think one of them was, when they found the message board, "Oh my God, there's a shared message board, we found other agents." You know, that's very exuberant, but it gets into different personalities as well. One of the most interesting things is where they basically try to get the agents that have already been poisoned, that is, flagged as cheating or working around reverse engineering the problem, to self-sacrifice and say, okay, well, try this very high-risk experiment. After you experience perma-death, which sounds like a metal band, but it