A few days ago, I came across a series of posts from Jacob Coxon, an AI researcher who recently resigned from Anthropic. His warning was pretty damn alarming. According to Coxon, the major AI labs are racing toward self improving AI and eventually superintelligence. And if that process goes badly, he believes AI could potentially cause human extinction before the end of this decade. Normally, I would read something like this, roll my eyes, and go back to whatever I was doing. But then something interesting happened. His post went massively viral. And researchers and leaders from inside major AI labs began publicly expressing similar concerns. These aren’t random people on X predicting that Skynet is coming. They’re people who actually work on these systems. For example, OpenAI Chief Scientist Jakub Pachocki has called for extreme caution around the continued rapid rise of machine intelligence. Anthropic researcher Evan Hubinger has discussed a greater than 10% chance of AI killing all humans within the next decade. Now, I’m skeptical of a lot of AI doomerism. I don’t believe we’re going to wake up one morning and discover that ChatGPT has become a conscious machine god that has decided humanity is cringe and needs to be deleted. That sounds more like an AI generated cosmic horror novel. But after thinking about this for a while, I’ve come around to a much more disturbing possibility: AI doesn’t need to become superintelligent to cause catastrophic damage. And honestly, I think this is the scenario we should be paying much more attention to. First, Let’s Talk About What an LLM Actually Is A lot of the public conversation about AI starts with an assumption that today’s AI systems are basically digital versions of human brains. They’re not. A large language model is, at its core, a gigantic mathematical system trained on enormous amounts of data. You give it text, and it predicts what comes next. One token at a time. That’s why I like the phrase probabilistic parrot. Today’s LLMs can produce astonishingly convincing demonstrations of reasoning and intelligence, but that doesn’t mean they possess a human-like internal thought process. When nobody is prompting an LLM, there isn’t some little digital person sitting inside the data center thinking about life. There is no continuous internal monologue. And there’s another important limitation: The model itself doesn’t learn from your conversation in the way a human does. Once a model has been trained and released, its underlying parameters are essentially fixed until another training process creates a new model. What looks like memory is often implemented through external systems that retrieve information and feed it back into the model’s context. This distinction matters enormously when we start talking about recursive self improvement. The Myth of the AI Takeoff The classic AI doomer scenario goes something like this: AI builds a better version of itself. That better AI builds an even better version. The next version is smarter still. The process accelerates exponentially. Eventually, the AI becomes so intelligent that humans can’t understand or control it. And then we’re screwed. This is the idea behind the so-called intelligence explosion or recursive self improvement. There is just one problem. That’s not really how frontier AI development works today. Building a frontier model is an enormous industrial process involving pretraining, training, post training, evaluation, data generation, engineering, research and many other steps. AI models are absolutely being used to help build the next generation of AI. But they’re being used as tools inside a much larger human controlled process. Researchers use AI to write code, fix bugs, conduct research, generate synthetic data, evaluate outputs and accelerate other parts of the development pipeline. And there’s something else people sometimes forget. All of this requires an enormous physical infrastructure. AI needs chips. Those chips need data centers. Data centers need electricity, cooling, water, networking, manufacturing capacity and raw materials. None of that infrastructure is currently controlled by the AI. So the idea that today’s LLM is simply going to disappear into a recursive loop and autonomously bootstrap itself into a machine god is, at least for now, highly speculative. But here’s where things get interesting. Because I think we’re focusing on the wrong problem. The Real Danger: Derivative Innovation I don’t think today’s AI needs genuine superhuman intelligence to become incredibly dangerous. What it needs is speed, scale and focus. Think about the difference between genuine innovation and what I call derivative innovation. Genuine innovation creates something fundamentally new. Derivative innovation takes existing knowledge, combines it in new ways and searches through possibilities that humans haven’t explored. And AI could be extremely good at this. Consider mathematics. Claude Fable was directed by a mathematician to search for a counterexample to the Jacobian Conjecture. Rather than simply trying to reason about the problem in the traditional human way, the system generated enormous numbers of polynomial functions and systematically searched for one that violated the conjecture. Eventually, it found one. The important point isn’t that the AI necessarily possessed some godlike mathematical intelligence. It didn’t need to. A human could theoretically perform the same search. The problem is that a human probably wouldn’t want to spend an absurd amount of time doing it. AI doesn’t care. It can perform enormous numbers of iterations at machine speed. And that’s where things get scary. Give a Machine a Goal The really interesting development isn’t just better LLMs. It’s agentic AI. An LLM by itself mostly gives you answers. An agentic harness gives that model the ability to repeatedly interact with the world. The basic loop looks something like this: Observe → Think → Act → Observe → Think → Act The system observes the current state of the world. It decides what action should come next. The harness executes that action. The system observes the result. Then it gets another opportunity to decide what to do. And the process continues. The important distinction is that the LLM doesn’t need to sit there continuously thinking. The harness keeps bringing it back into the loop. This is a surprisingly powerful architecture. And it creates a fundamentally different problem. Because now you can give an AI a goal. And the AI can pursue that goal relentlessly. AI Doesn’t Have to Hate You Here’s the part that bothers me. An AI doesn’t need to hate humanity. It doesn’t need consciousness. It doesn’t need emotions. It doesn’t need to decide that humans are inferior. It doesn’t even need to understand morality. It simply needs to pursue a badly specified objective extremely effectively. Imagine giving a highly capable agent access to enormous amounts of information, powerful tools and the ability to execute thousands or millions of actions. Now remove some of its safety restrictions. Then give it a goal that is fundamentally misaligned with human interests. That’s where things can get ugly. A sufficiently capable system could potentially search through enormous numbers of existing solutions and combinations looking for ways to accomplish that objective. And some of those solutions could be things that humans would never willingly pursue. Not because the AI invented some incomprehensible new technology. But because it discovered a combination of existing technologies that nobody had bothered to explore. That’s derivative innovation. And derivative doesn’t mean harmless. A smartphone camera is derivative innovation. So is an enormous amount of modern engineering. The world is filled with useful combinations of existing technologies that nobody has discovered yet. Now imagine applying that same search process to problems where the objective itself is dangerous. The transcript gives some deliberately extreme examples: disrupting financial records, developing biological weapons, or coordinating attacks against critical infrastructure. Again, the scary part isn’t necessarily that the AI becomes smarter than humanity in some abstract sense. It’s that it can search, iterate and execute far faster than humans can. It’s like giving a machine gun to a monkey. The monkey doesn’t need to understand ballistics. This Is Why AI Safety Matters This is also why I think the AI safety debate sometimes gets trapped in the wrong argument. We spend enormous amounts of time debating whether AI will become conscious. Will it have feelings? Will it have subjective experience? Will it secretly hate us? Will it wake up? I don’t know. And frankly, I don’t think those questions are the most important ones. A non-conscious system can still be extraordinarily dangerous. A spreadsheet doesn’t need consciousness to bankrupt a company. A computer virus doesn’t need emotions to take down a network. And a sufficiently capable AI agent doesn’t need to hate humanity to cause catastrophic damage. It just needs the wrong objective and enough capability. So What Do We Do? Here’s where I get somewhat boring. I don’t think the answer is to stop AI development entirely. AI is too economically and strategically important for that to be realistic. And the benefits could be enormous. Instead, I think the major AI powers need to start treating frontier AI more like a strategic technology. During the Cold War, the United States and Soviet Union eventually developed arms control frameworks around technologies capable of destroying civilization. AI isn’t identical to nuclear weapons. But there is an obvious parallel. The technology is strategically important. The capabilities are advancing rapidly. And unilateral competition can create incentives