Dr. Alex Turner went viral for resigning from Google DeepMind over its Pentagon AI contract, which he said lacked binding restrictions against autonomous weapons and mass surveillance. Alex is a world-class researcher with a PhD in alignment from Oregon State University, a postdoc at Stuart Russell’s Center for Human-Compatible AI at UC Berkeley, and top distinctions at NeurIPS. He’s one of the minds behind Shard Theory (with Quintin Pope), and he also helped pioneer the Steering Vectors research made famous through Anthropic’s Golden Gate Claude experiment. Alex says Google DeepMind broke its founding promise — the commitment Google made when it acquired DeepMind — and that the people best known inside Google for caring about AI ethics largely didn’t act when it counted. Alex came up through the rationalist community, and became one of LessWrong’s highest-karma users before leaving the site. He still shares many of my concerns about AI doom, but he believes technical alignment research is going better than expected. I’m not so optimistic. In this episode, we debate my Yudkowskian doom views against Alex’s own framework. Can he convince me the Yudkowskians are miscalibrated? P.S. We’re currently in a donation drive, so here comes our solicitation for viewer donations… 👉 Please click here to make a tax-deductible donation to Doom Debates 👈Doom Debates is a fiscally sponsored project of Manifund.org, a registered 501(c)(3) nonprofit. To donate crypto or if you have questions, email me. Watch on YouTube Timestamps 00:00:00 — Cold Open00:01:15 — Introducing Alex Turner00:02:30 — From Harry Potter Fanfic to AI Alignment00:05:03 — Meeting Quintin Pope & Rethinking AI Doom00:06:17 — Shard Theory, Steering Vectors & Golden Gate Claude00:08:25 — Why He Joined Google DeepMind00:10:32 — Google DeepMind’s Broken Promise00:16:01 — Debating Google DeepMind’s Pentagon Contract00:19:09 — What’s Your P(Doom)™?00:22:42 — Alex’s Research on Instrumental Convergence00:26:35 — Misuse vs. Misalignment: The Mainline Doom Scenario00:29:31 — Will Society Self-Correct?00:36:36 — Superintelligence in 10 Years00:39:43 — Will Technical Alignment Produce a Safe AI?00:41:16 — Donation Drive00:42:12 — How Fragile Is the Chain of Alignment?00:49:50 — Disagreements with Yudkowsky’s ‘List of Lethalities’00:57:46 — Why Alex Quit LessWrong01:01:48 — What’s Next for Alex01:02:54 — Does He Support PauseAI? Stop the AI Race?01:04:40 — Wrap-Up01:06:04 — Producer Ori’s Closing Note Links Alex Turner, “Why I Left Google DeepMind” blog post — https://turntrout.com/why-i-left-google-deepmind Alex Turner (TurnTrout), personal website — https://turntrout.com Alex Turner’s resignation announcement on X — Doom Debates episode with Quintin Pope — https://lironshapira.substack.com/p/ai-alignment-is-solved-phd-researcher Harry Potter and the Methods of Rationality (HPMOR) — https://hpmor.com/ Shard theory sequence on LessWrong —https://www.lesswrong.com/s/nyEFg3AuJpdAozmoX Golden Gate Claude (Anthropic research on steering vectors) — https://www.anthropic.com/research/golden-gate-claude Slaughterbots on YouTube — Alex Turner, “Avoiding Power Seeking by Artificial Intelligence” PhD thesis — https://turntrout.com/alignment-phd Alex Turner, “Some of My Disagreements with List of Lethalities” — https://turntrout.com/disagreements-with-list-of-lethalities Doom Debates episode with Bentham’s Bulldog — https://lironshapira.substack.com/p/benthams-bulldog-ai-doom-debate Alex Turner on X — https://x.com/Turn_Trout Doom Debates donation page — https://doomdebates.com/donate Transcript Cold Open Liron Shapira 00:00:00You resigned from Google DeepMind because you think that Google DeepMind, quote, “Broke its founding promise through its contract with the US military.” Alex Turner 00:00:08I value staying true to your values. People who are well-known within Google for caring about the ethics of deploying AI largely didn’t act. This matters because autonomous weapons get us into an arms race that really degrades the security of everyone in the world. Liron 00:00:26Do you support the Pause AI movement? Alex 00:00:28I think that AI is being developed too quickly. I probably will not take an affirmative on supporting this particular movement. Liron 00:00:35Let’s segue into the schism, your disagreement with Eliezer Yudkowsky. Alex 00:00:39He was incorrect on some points for alignment, but then also not acknowledging that. I think that technical alignment has gone pretty awesome, super awesome. Liron 00:00:49You’re not claiming that a superintelligent AI can’t kill everybody. You’re like, “Oh yeah, of course it can, but we’re not gonna break the chain of alignment, meaning we’re just going to safely develop it so that even though it can kill everybody, it won’t.” Alex 00:01:00Develop it in a way that produces a safe result. I wouldn’t call what we’re doing safe development, but sure. Introducing Alex Turner Liron 00:01:15Welcome to Doom Debates. My guest has worked in technical AI safety at Google DeepMind for two and a half years, but he just quit, and his resignation is going viral. Why? Alex Turner says Google DeepMind, quote, “Broke its founding promise through its contract with the US military.” His latest blog post exposes hypocrisy at the highest levels of senior leadership at Google DeepMind, which includes the CEO Demis Hassabis, chief scientist Jeff Dean, and co-founder Shane Legg, among others. Alex is a world-class AI alignment researcher. He completed a PhD in alignment from Oregon State University. He did a postdoc at UC Berkeley, and he’s earned top distinctions at NeurIPS. I respect that Alex is principled. I respect his mastery of the subject matter that we talk about on this show. I also find it interesting that he’s levied criticism at the original AI alignment thinker, Eliezer Yudkowsky. He’s called some of Yud’s claims fundamentally misguided, not reasonable, and bogus. As a Yudkowskian myself, I’m gonna be curious to dig into those arguments. And of course, we’ll cover what’s going on right now at Google DeepMind and why he resigned. Alex Turner, welcome to Doom Debates. Alex 00:02:28Hey, thank you for having me. From Harry Potter Fanfic to AI Alignment Liron 00:02:30So it’s great to get you on the show. One thing we do on Doom Debates is we expose top intellectuals who have been pretty familiar to the rationality community or the AI safety community, and we help popularize their ideas, even if I don’t fully agree with all of them. Is that a good description of your background? You’ve been pretty deep into the LessWrong rationality and alignment community for a while. Alex 00:02:52Yeah, I think it was quite formative. It was just the other day in 2016 where I decided to search what are the top five Harry Potter fan fictions. And that indeed led me down the LessWrong rabbit hole, where I discovered superintelligence in late 2017, and then I pivoted my PhD in early 2018. From that time period up through maybe early 2023, LessWrong was very central to my professional career, but also just to the way I looked at the world. Liron 00:03:25Well, I wanna follow up on why did you search for Harry Potter fan fictions? Alex 00:03:30I really don’t know. It’s kind of one of those things where if I hadn’t done it, my life would be totally different. The reason I’m mentioning this is there’s this famous fan fiction that Eliezer wrote called Harry Potter and the Methods of Rationality. I never really liked fan fiction. I thought it was kind of cringe. No one recommended it to me. So it seems to me like if I’d just woken up slightly differently that morning, I might not have ever been exposed to this research area, and my life would be totally different. Liron 00:04:02Wow. And you said this is all in 2016, right? So the book had been mostly completed at that time. It had been going on from 2009 to 2015, and you kind of stumbled on it because you were just interested in seeing what the best Harry Potter fan fiction was? Alex 00:04:17I had a random thought. That’s the best I recall. Liron 00:04:21It’s pretty crazy that that’s how you found the community because I know you as one of the highest karma LessWrong users. You have ten times my karma. You’ve been posting a lot. It became a huge passion for you, right? Alex 00:04:32LessWrong was for quite a while my intellectual community. I’d have ideas. I’d be eager to share them. Each summer I would generally do an internship where I’d come in person at Berkeley, get to hang out with my friends there, be able to — I guess I felt more understood. In 2018, 2019, 2020, 2021, these are many years where I was at my PhD and I talked about the dangers of AI and how we should work on that. Meeting Quintin Pope & Rethinking AI Doom Liron 00:05:03Did you originally feel like you bought into all the Eliezer Yudkowsky concepts, and then you started rethinking everything and building it from the ground up? Was there a point of divergence? Alex 00:05:15Yeah. I think I was maybe around 80, 85% doom conditional on developing AGI. I thought it’d be a couple decades, even late 2021. And yeah, I shared most of the worldview. I found much of his writing compelling, and I still think there’s some gems in there. It wasn’t until early 2022. I met a researcher at my university, at Oregon State University, named Quintin Pope. He wrote these very big brain Google Docs, and he was sending them by me. He’d attended my AI alignment reading group. I don’t know, something was just very interesting about them, and they seemed really far-fetched, but they were very ambitious. And as I looked more, I realized he was pointing out some real confusions, real issues. I started rethinking perhaps the claimed difficulty of alignment. Sh