THE NEXT BILLION SECONDS: AI@70 EP 2: THE WATERSHED ChatGPT opened the door to artificial intelligence. Almost immediately, researchers built 'agents' using it - but three years of hard yards passed before agents were 'good enough' to be useful. In this episode we trace the path from 'bad' to 'good enough' agents - and what it unlocked as we 'crossed the watershed'. No one really thought that on the last day of November in 2022, the world would change completely. Not even the folks closest to it. Like Sam Altman, CEO of OpenAI, who said, "We always knew we'd hit a tipping point," "But we didn't know what the moment would be." His firm modestly promoted a 'research preview' of their new AI chatbot.ChatGPT. G'day, I'm Mark Pesce and this billion seconds are already unfolding as the most significant of this century. In this miniseries, celebrating the 70th anniversary of artificial intelligence, we're looking at where we’ve come from, how we got here - and where we seem to be going. Because we’re travelling at the speed of thought. ChatGPT promised artificial intelligence on tap for everyone. But it took another three years for that promise to be realised. This episode traces that story, from something kids used to do their homework for them - into a tool reshaping our economy. That's on this episode of This Billion Seconds. I reckon anyone who used a chatbot in 2023 or 2024 deserves a gold star. Two gold stars if they used one to get work done. ChatGPT was interesting. But it wasn't very accurate. The earliest chatbots, trained to be eager to please, would sometimes make up their answers. With no basis in fact. We call these 'hallucinations', but that's not what they are. It's that we'd caught the AI out. Asked a question it didn't know how to answer. So, it did its best with whatever it could pull together. It didn't mean to lie. It didn't even know it was lying. It was just generating a response. In a chatbot conversation that isn't a big deal. You could fact check - against another chatbot, against the web - even, if you can imagine it, ask another human being. Those sorts of hallucinations could be caught - if you went to some effort. But then again - why ask a chatbot if you want to make an effort? Folks took chatbots at their word. And some them paid a price. Hallucinations are an annoyance for people. But they're show-stoppers for autonomous agents. Oh yes, the nearly forty year old dream of Apple CEO John Sculley of autonomous agents running around and doing all of our work for us - that dream flickered back into life alongside ChatGPT. Researchers reckoned that with 'good enough' AI, they'd be able to build agents. To read your email, Keep your calendar. Book a reservation. That sort of thing. Just three months after ChatGPT landed, the first of those autonomous agents popped up. A piece of software known as AutoGPT. AutoGPT turns an AI like ChatGPT into an agent. It does that by providing the three things an agent needs. Memory - so that the agent can remember what it's supposed to be doing, and keep notes on its progress as it does it. Tools - So that it can do things like read and write files, respond to emails, or add items to the calendar, And goal logic - this is the thing that turns an AI into a single-minded goal-oriented piece of software. Basically, it's the same quality as the Terminator: the agent has one goal and it will stop at nothing to achieve it. AutoGPT gave ChatGPT memory, tools and goal logic - everything need to turn it into an autonomous agent. And it should have been absolutely amazing. Except for one thing. Here's how an agent works: you give it a goal, and it 'decomposes' that goal into a series of steps, then breaks those steps down into discrete actions. The agent then methodically works its way through each of the actions. At the end of every action, an agent 'reflects' - it checks the results of that action. Did it work? Does the agent get to go on to the next action, or does it need to do this action again? And this is where the problems arise. Because if ChatGPT happens to hallucinate in the midst of this process, the whole thing very quickly falls over. Maybe an agent thinks an action worked, when it didn't. Or thinks it didn't, even though it did. Now there's always a change that - for any given action - there will be a hallucination that will make the agent fall over. You can deal with that if there are just a few actions. Chances are it will get through all of the actions before problems arise. But for anything even modestly complex, there are many actions. Tens to hundreds. Maybe even thousands. So the probability for just one hallucination - that's all it takes to make the agent fall over - grows higher and higher as the list of actions grows longer and longer. Now here's the thing: people measured this. A group known as Model Evaluation and Threat Research or METR, they've kept a running record of how long agents can run before they fall over. And back in early 2023, they'd make it no more than about 4 minutes. An agent can't do a lot in 4 minutes. So although we could make agents using ChatGPT, we couldn't make them work well. For that, we'd need better AI. Funny thing about that. All of us can take some credit for making AI better. You see, most everyone who's using ChatGPT or Claude or Gemini or any of the others, is making those models better with every question we put the them. Those questions and the answers to them get fed back in, to train the next generation of models. It's why AI is so much better now than it was three years ago - and why I say anyone who used AI day-to-day in 2023 or 2024 deserves a gold star. Those were hard yards because the AI just wasn't very good. How do we know AI is better to day than it was two or three years ago? Ah, here's where we come back to those folks at METR. They've been tracking what they call the 'task horizon' of agents. How long they can perform actions before they fall over. And they noticed that with every subsequent generation of AI, that task horizon got longer. Dependably. It soon became clear that the task horizon doubled an average in seven months. 4 minutes becomes eight minutes toward the end of 2023, then sixteen minutes in mid 2024 thirty-two minutes in early 2025, and then sixty-four minutes. Just over an hour In November 2025. Three years after ChatGPT launched. Now that moment in time - just under a year ago - saw the launch of three brand new models: Google Gemini 3, Anthropic Claude Opus 4.5, and OpenAI GPT-5.2. Each of these models could be used to create agents with task horizons greater than an hour. An hour is "long enough" that you can assign an agent a reasonably complex task, let it go off and do the work, with confidence that the work will be done correctly. That's kind of a magic length of time. It takes agents out of the theoretical - where they'd been for nearly 40 years - and makes them very practical. That's the reason I call this moment "The Watershed". It's the moment when AI gets "good enough" to do real work. I'm far from the only person to notice this. Anyone using AI tools to write software noticed a big shift, as they stopped fighting with their tools. Because the tools had gotten 'smart enough' to handle the task. This is the moment where we first hear the term 'vibe coding' - tell the agent what kind of software you want to create, and the agent will go off and build it for you. And now that we're on the other side of the watershed, it's all downhill. Gaining speed. Because the task horizon didn't stop growing at an hour. We're more than seven months beyond November 2025, and task horizons have doubled again. Two hours. By early next year, four hours. And by the end of next year - eight hours. An entire work day. Tell the agent what to do - for the whole of the day - and let it do its thing. We've come a long, long way since November 2023. But we're through the hardest bits. You're using the worst AI you'll ever use. And boy, will it get better from here. In our next episode, we'll look at what autonomous agents mean for business - and whether business is actually prepared to pay for them. That's on the next episode of This Billion seconds. THIS BILLION SECONDS was written and recorded by Mark Pesce. Produced with assistance from Myrtle and Pine. If you like this show, please share it with a friend. And make sure to follow or subscribe to get all of the episodes in this series.This is Mark Pesce, thanking you for listening. See omnystudio.com/listener for privacy information.