This episode is with Jean-Stanislas “JS” Denain of Epoch AI, who leads their Insights Team and is one of the people I find myself debating the state and trajectory of AI with more and more. We’ve had follow-on discussions of many of my favorite recent posts online and/or in private, so I wanted to dig into the nuance in a public episode. A big takeaway of this podcast is how JS and I both have so much uncertainty with exactly where we are heading, and this was our best effort at stating our observations today. Chapters / topics include: * 00:00 Predictions for RSI * 18:15 The role of robotics in an AI acceleration * 24:20 How far behind are Chinese models? * 27:39 Does distillation explain the gap? * 40:58 What Chinese job postings reveal about their labs * 48:13 Are open or closed models safer? * 58:10 How Epoch AI ticks * 1:00:55 What a frontier post-training recipe looks like Enjoy! More from JS: Epoch AI profile and writing, X, LinkedIn Listen on Apple Podcasts, Spotify, and where ever you get your podcasts. For other Interconnects interviews, go here. Transcript 00:00:00 Nathan Lambert: I’m here with JS Denain, who is a senior researcher at Epoch AI. He leads the insights team. He is one of the people who I feel like I get the best feedback on my writing from, whether it’s from US-China AI capabilities, now RSI. And I just wanted to open this discussion and honestly go deeper with him, trying to understand how he thinks about these various things. And I think you have a very useful, moderate point of view, which I feel like you’re probably a step further into what would be called faster scenarios for AI progress. But let’s get into this, and it’s like, what measurements do you think OpenAI and Anthropic are seeing when we get all these proclamations on RSI happening very imminently? 00:00:47 JS Denain: Yeah. I think, so there’s the measurements they’ve published, right? So, OpenAI and Anthropic both had blog posts, I mean, Anthropic two at least, on the effect AI has on accelerating AI progress. I think at least the things they publish, I don’t think are super strong evidence of imminent self-sustaining acceleration AI capabilities, or full automation of the job of AI researcher. But I think the kinds of things we see are, I think probably the most striking thing I saw in the OpenAI blog post was increasing usage of AI systems in model deployment, like the increase in spending on Codex that we saw. And it’s kind of unclear how exactly to interpret this, because maybe it’s a measurement artifact where they’re only looking at Codex, but in fact, there was a bunch of ChatGPT usage before from the researchers. But overall, that plot, for example, just shows a 2X a month increase in Codex spending by researchers, and that does seem to me to be some evidence of they’re getting a lot of value out of this probably. I don’t think this is strong evidence that in six months we have a software intelligence explosion. 00:01:57 Nathan Lambert: Do you think this is the same? So, what is the information they have internally relative to what we have? And this is obviously hypothetical. We don’t have this internal information. Because I get the sense that a lot of people are more scared in their updates from the labs than the information we have. And I try to take this very seriously of, what will they be seeing that is making the acceleration of risk comments go faster, and how much of this is material evidence versus how cultures evolve over time? And I’m much more interested in evidence. 00:02:29 JS Denain: Yeah. So two things. I think, first of all, I guess I don’t think, I don’t know, right? I don’t have full information here. I don’t currently think that either there’s some specific thing that people at OpenAI or Anthropic are seeing right now that we don’t have access to that warrants being way more freaked out about this. I also don’t think that... I think the public evidence we have right now, more general on AI progress and just a priori case for this being an important dynamic, I think is enough. I think to care about this particular dynamic of AI accelerating AI progress, that being a big deal and worth tracking. And then it’s kind of unclear what the urgency is of when the feedback loop really kicks in. So, basically on the what is there on the inside that people have access to, I could give examples of kinds of metrics, right, they could be looking at. It’s plausible that we have access to the capabilities of AI systems, but internal teams have their KPIs, and maybe they’re seeing compute multipliers in the pre-training team or other kinds of metrics that people are tracking going crazy. And then the combination of this plus some intuitions of how the different outputs of different teams combine yields a prediction on the trend in actual performance of the end AI systems. So that could be an early warning sign. It’s unclear to me that the recent discourse we’ve seen is evidence of things going crazy on those metrics. 00:04:09 Nathan Lambert: And how do you think of the link between RSI and existential risk? So I would posit that you agree. I think that there are very real risks of AI, and I’m curious on how you think these, what I would describe as very, very early measurements change anything on the scope of risk. Because I don’t think if you had asked people six months ago, it would be as immediate to x-risk among people are very reasonable. I think there’s more people that are reasonable talking about x-risk again, which was a little surprising to me. 00:04:43 JS Denain: Yeah. So, okay, my sense is something like... So personally, I feel very uncertain about this, but I do feel, yeah, basically bought into there’s, I know Evan Hubinger was like, at least 10% of x-risk within I don’t know what timeframe. I think I’m like, yeah, I know, and this seems pretty reasonable over a decade-long timeframe. I just feel extremely uncertain about it, but I’m definitely very worried about this. Now, why am I worried about this, and where do I think the disagreements come from? And then how do I relate this to the early sense of RSI? My sense, I’m kind of a capabilities theory of everything person. I think, and I think some people disagree here, but I really think that principal component of disagreement between everyone is how huge do the capabilities get, how soon, of AI systems? And I sort of agree that there’s other factors that come in play for how big your economic growth gets, also depend on the diffusion you get. And you could possibly you could think that capabilities are going to get crazy, but the AI system’s going to be just aligned and benign and stuff, and so there’s no huge risk. But my sense is concretely, when I look at the main kinds of disagreements between people, most of the people who I see who are very skeptical of those most extreme scenarios... I think just expect capabilities to not be as huge or as I think folks who are— 00:06:18 Nathan Lambert: What does being a capabilities maximalist look like in a few years? Because I think I’m probably on the skeptical side, so please continue. 00:06:28 JS Denain: I think it looks, for example, something like the AI 2027 scenario, right? I think it looks like the mechanism for this is AI is automating the AI research process, I think, and that’s a reason to pay attention to it. But in terms of effect on the real world, I think it’s like massive progress on robotics. I think a huge industrial explosion, AI systems are just managing factories. You have this kind of self-sustaining economy that just is able to make a large scientific progress much faster than you would have expected. And so I think concretely, the kinds of disagreements I would expect are on, yeah, if you have AI systems that are both very intelligent in the book smart sense, but also have been trained to have more affordances and use them astutely, have been trained to kind of manage projects in efficient ways and stuff like that. How big are the real-world bottlenecks to making very fast R&D progress or getting hard power over humans? 00:07:26 Nathan Lambert: Yeah. Can we go into some of these in detail? Did you listen to the Dwarkesh podcast with Charlie, Beren, and John? 00:07:32 JS Denain: Yes. Yeah. 00:07:33 Nathan Lambert: Yeah, because they had at the end, they had this section on various capability levels and timelines for getting them. And I feel like I agreed with most... I was very in agreement on the distribution they had up to this, and then was surprised by the timelines. And one of them was the 10X productivity for the AI researchers. And I think Beren and John were faster than I think. And my kind of statement is that I think the cycle from of having an idea and doing the experimentation to test it, I agree will be 10X faster very soon. But I don’t necessarily agree that I would say that AI researchers will be 10X more productive in net, which I would describe as the pace of the field’s complete understanding. And understanding is a different axis from just continuing to scale models. I think that’s one of my core confusions on the AI research side. So I’m just kind of curious how you think about this type of thing and how you might specify a 10X improvement in AI research into subcategories. 00:08:37 JS Denain: Yeah. So maybe there’s a scale you could have here, which is the end thing that you might care about is how much faster is AI research overall? Or how much faster is Anthropic’s overall output? And then Anthropic as a company is producing some things, and it’s doing in one year what it would have taken it 10 years to do. And that’s pretty different from individual researcher productivities, where I think you could... So if most of what AI researchers right now are doing is this loop that you were describing, then it’s possible that you get a 10X productivit