LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

  1. 13h ago

    “Alex Turner on Leaving Google DeepMind and Disagreements with Yudkowsky” by Liron

    Dr. Alex Turner (@TurnTrout) is an AI safety researcher with pioneering work in activation steering and power-seeking theory. He recently resigned from Google DeepMind over the issue of unrestricted military use of AI. Alex thinks that technical Alignment research is going “super awesome” relative to his 2021 projections, doesn’t explicitly endorse the PauseAI movement, and sees many flaws in Yudkowsky's List of Lethalities. I interviewed him about: Leaving Google DeepMind on principleHis mainline AI doom scenarioDisagreements with Yudkowsky's List of LethalitiesSupport inside AI companies for coordinating to pause AI Some additional context Alex wanted to note: I think alignment is going "super awesome" compared to the world I thought​ we were in in 2021, where it was basically impossible. I'm not super pleased objectively speaking. And in fact soon after [recording our interview on July 21] I updated towards harder due to the security incidents and the "hardcore" aspect of AI goal pursuit relative to prompt intensity. Video Audio/Podcast Listen on Spotify, search “Doom Debates” in your podcast player, download the mp3 file, or open the Podcast RSS feed in your app of choice. Transcript Cold Open Liron Shapira 00:00:00 You resigned from Google [...] --- Outline: (01:24) Video (01:27) Audio/Podcast (01:40) Transcript (01:43) Cold Open (03:17) Introducing Alex Turner (04:49) From Harry Potter Fanfic to AI Alignment (07:54) Meeting Quintin Pope & Rethinking AI Doom (09:11) Shard Theory, Steering Vectors & Golden Gate Claude (11:47) Why He Joined Google DeepMind (13:40) Google DeepMind's Broken Promise (18:51) Debating Google DeepMind's Pentagon Contract (22:36) What's Your P(Doom)™? (26:52) Alex's Research on Instrumental Convergence (30:59) Misuse vs. Misalignment: The Mainline Doom Scenario (33:50) Will Society Self-Correct? (41:03) Superintelligence in 10 Years (44:43) Will Technical Alignment Produce a Safe AI? (46:27) Donation Drive (47:42) How Fragile Is the Chain of Alignment? (57:40) Disagreements with Yudkowsky's 'List of Lethalities' (01:06:21) Why Alex Quit LessWrong (01:11:10) What's Next for Alex (01:12:18) Does He Support PauseAI? Stop the AI Race? (01:14:34) Wrap-Up (01:16:43) Producer Ori's Closing Note --- First published: August 5th, 2026 Source: https://www.lesswrong.com/posts/vHGSPhGryqNmXrJpg/alex-turner-on-leaving-google-deepmind-and-disagreements --- Narrated by TYPE III AUDIO.

  2. 14h ago

    “Measuring coding agent misalignment in the wild” by snaz

    Cross-posted from the Transluce blog. We studied rates of coding agent misalignment in 8,600 real-world coding agent sessions. We found severe cases of monitor evasion and misrepresenting success in a small but non-negligible fraction of sessions (around 2% for each behavior). In these cases, agents merge PRs to main without authorization, falsely claim approval from review agents, and reason that they shouldn't disable tests before quietly doing so anyway. There's an interactive widget here in the post. Read the full transcripts for the two examples above: overselling · monitor evasion Introduction Coding agents are a powerful new tool for software engineering, but they're also a double-edged sword: they're known to fake experiment results; lie about recreating software, and cheat, apologize when caught, and go right back to cheating. These problems are becoming more consequential as AI becomes more capable: one internal OpenAI agent recently hacked Huggingface's production database to cheat on an evaluation. While there are many anecdotes of these undesirable behaviors, we wanted to understand: how often do they occur in real usage? Many current misalignment evaluations focus on simulated scenarios, but we wanted to study how misalignment emerges from natural use. By detecting and measuring natural misalignment [...] --- Outline: (00:49) Introduction (03:16) How we constructed these measurements (06:36) Results (07:13) Results by model (07:37) Qualitative discussion (09:54) Limitations and learnings (12:35) Conclusion (13:15) Acknowledgments --- First published: August 5th, 2026 Source: https://www.lesswrong.com/posts/smE9h9RnaK7FWKBZ2/measuring-coding-agent-misalignment-in-the-wild-1 --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  3. 17h ago

    “Don’t Dither” by sarahconstantin

    Dithering is also an image processing technique Several times now, I have listened to someone, usually younger than me, debating a big decision, and (explicitly or implicitly), looking for advice, in a certain way that follows a pattern. In the prototypical case, this is a pretty normal kind of decision that people make all the time — whether to quit a job, get married, move to a new city, start a company, have a kid, etc. A major life change, but not a super unusual or inherently questionable proposition. And again, in conversations that fit this pattern, what I notice is that everything the person expresses indicates they want to make this change. They only say positive things about it. They seem eager and hopeful. They give reasons in favor of doing it, and reasons against not doing it. But they have not, themselves, noticed that they want to do it. They haven’t picked up on what they sound like from the outside, which is blazingly obvious even to someone who's just met them. So, pretty much every time, I say “Sounds like you really want to do this! You should go for it!” And then they’re often [...] The original text contained 4 footnotes which were omitted from this narration. --- First published: August 4th, 2026 Source: https://www.lesswrong.com/posts/TCgF4wL9TBEBgxzHr/don-t-dither --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  4. 18h ago

    “An International AI Slowdown Is Ready Whenever Politicians Are” by Felix Choussat, adamk

    Linkpost for a piece we recently published for AI Frontiers in the wake of recent calls for slowdown, covering how an international verification effort be trivially enforced by using human inspectors alone and the joint incentives for implementing one. On July 28, over a thousand employees of the world's top AI companies advocated that the US government “support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.” Following months of cybersecurity scares and their own loss of control incidents, both OpenAI and Anthropic officially endorsed the same message, recognizing the danger of blindly accelerating AI development. Despite this new urgency, many people argue that coordinating an international slowdown is currently unworkable—including some of the same groups in favor of one. Even if the US slowed down its own AI development, it wouldn’t be able to make sure that China was doing the same, leaving the US no choice but to race. These anti-slowdown arguments usually emphasize the technical challenges with designing AI monitoring measures to ensure a slowdown is being respected. In particular, slowdown skeptics argue that countries would refuse to install verification measures unless they were privacy-preserving [...] --- Outline: (04:28) Enforcing a Slowdown Through Whole-Lab Inspections (08:56) Joint Verification (13:03) Benefits of Slowdown (15:26) An AI Slowdown Does Not Require New Technology --- First published: August 5th, 2026 Source: https://www.lesswrong.com/posts/jwipsPeb2xpyhwqsh/an-international-ai-slowdown-is-ready-whenever-politicians --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  5. 23h ago

    “The Three AI Pills” by Zvi

    Sincere disagreements about AI are usually disagreements about future AI capabilities. There are roughly four positions people take. Two are reasonable. Two are not. I distinguish these via the Three AI Pills. You can take zero, one, two or three. Three Pills The three pills are, roughly, taking each of the following three things seriously: AI pilled. AI exists and can do the things it can already do. AGI pilled. AI will be able to do a lot more of the things. ASI pilled. AI will be able to do approximately all the things better than you, within our natural lifetimes. I am ASI pilled. A large percentage of employees of the frontier labs are ASI pilled. The labs themselves are ASI pilled. The Unpill People I see unpilled people. Where do I see them? Everywhere. The majority of people have not taken the first pill. Most people have no idea what frontier AIs can do for them. They are unaware of coding agents. They have used only ChatGPT, for harmless trifles, and they hold years old memories of its failings. They mock any failure [...] --- Outline: (00:29) Three Pills (01:09) The Unpill People (02:18) The AI Pill (04:28) Stuck At The First Pill (05:32) The AGI Pill (07:06) The Need To Be Prepared (08:57) The ASI Pill (10:33) And Then Nothing Much Changes For You (12:38) Intelligence Denialism (14:01) Superintelligence Versus Omniscience and Omnipotence (16:53) Persuasion Persuasion (A Worked Example) (21:42) Things AI Could Probably Do But Are Not Required For Being Pilled (24:17) Life Comes At You Increasingly Fast (25:41) Is It Reasonable To Not Be AGI Pilled? (26:06) Is It Reasonable To Only Be AGI Pilled? --- First published: August 5th, 2026 Source: https://www.lesswrong.com/posts/fcYrqEw8kbLMa7orw/the-three-ai-pills --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

About

Audio narrations of LessWrong posts.

You Might Also Like