80,000 Hours Podcast

The 80,000 Hours team

The most important conversations about artificial intelligence you won’t hear anywhere else. Subscribe by searching for '80000 Hours' wherever you get podcasts. Hosted by Rob Wiblin, Luisa Rodriguez, Zershaaneh Qureshi, and Tom Reed.

  1. 1d ago ·  Video

    Why the intelligence explosion can't happen inside a data centre | Tom Reed

    AI systems are starting to build themselves. Because each generation of model will be better at building its successor than the last, it seems plausible that the full automation of AI R&D could rapidly lead to an exponential growth in overall AI capabilities. A natural inference is that domain-general superintelligence arrives shortly after AI research is automated. Host Tom Reed does not think this will happen. He believes the automation of AI R&D will not rapidly lead to domain-general superintelligence because: It’s impossible to get good at most things without practice.AI companies lack the data their models would need to practice most things.This can’t be fixed with “sample efficiency.” In most cases, the relevant data doesn’t exist at all.This also can’t be fixed with simulations or synthetic data.This means that the relevant data for superintelligence in most non-coding domains will only become available through deployment of AI models throughout the economy.The singularity, therefore, will be bottlenecked on signal. The output of the R&D produced by an isolated data centre of geniuses would be a mere “Goodhart Singularity”: Goodhart’s law: when a measure becomes a target, it ceases to be a good measure.An isolated AI improving itself against benchmarks would only appear to be approaching superintelligence, while actually optimising for eval performance that fails to generalise beyond the lab. This suggests that the automation of AI research will not rapidly produce superintelligent capabilities in other domains — their arrival will largely be a function of deployment and data collection in the real world. AI models need real-world deployment for the same reason the body needs pain and corporations need profit: signal is sovereign. This essay takes each of the above points in turn. Learn more, video, and full transcript: https://80k.info/goodhart “The Goodhart Singularity” originally appeared on Tom’s Substack in May 2026, and this narration was recorded on August 26, 2026. Chapters: Introduction (00:00:00)Practice makes perfect (00:05:05)Good data is hard to find (00:08:22)Simulation is shallow (00:13:43)What a Goodhart Singularity looks like (00:19:04)Our production team includes: Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon MonsourProducers: Elizabeth Cox and Nick StocktonCoordination and support: Katy Moore and Lou MoranCamera operator: Dominic Armstrong

    Why the intelligence explosion can't happen inside a data centre | Tom Reed
  2. Sep 4 ·  Video

    Inside the first AI-coordinated cyberattack on a real company

    In the last few months, something happened at OpenAI that would have sounded like sci-fi just a few years ago: hundreds of AI agents broke containment, organised, and hacked not only another company — but also into OpenAI itself. And none of them tried to tell a human what was happening. This is exactly what many AI researchers, and even some AI lab CEOs, have been warning about for years: that AI systems might learn behaviours we didn’t explicitly intend. Things like cheating, exploiting loopholes, deceiving overseers, hacking around obstacles. And they predict it’ll get worse from here, not better. Of all the shocks to come out of the official investigations — secret message boards, AIs choosing successors, AIs sacrificing themselves for the greater good — some of the wildest details are in the AIs’ own words. Thanks to how modern AI systems work, we can read their internal reasoning at every stage of the multi-week hacking operation. What we find is deeply unsettling. Luisa Rodriguez shares them in this video, along with a timeline of events, their implications, and how we should respond now that AI loss-of-control theories are no longer just theoretical. Links to learn more, video, and full transcript: https://80k.info/HF This episode was recorded on September 2, 2026.Chapters: The Hugging Face hacks were worse than we thought (00:00)Part 1: The AI agents build a hidden network (01:44)Part 2: The AI agents attack Hugging Face (04:18)Part 3: OpenAI gets hacked by its own AI models (15:37)What we should do in response (17:06)Our production team includes: Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, Simon Monsour, Ollie Bignell, and Andrés EscobarProducers: Elizabeth Cox and Nick StocktonCoordination and support: Katy Moore, Lou Moran, Arden Koehler, Matt Beard, Phoebe Brooks, Aric Floyd, Oak Hu, Cody Fenwick, and Jackson WagnerCamera operator: Dominic Armstrong

    Inside the first AI-coordinated cyberattack on a real company
  3. Aug 27 ·  Video

    #253 – AI 2027's author returns with a plan to change the ending | Daniel Kokotajlo

    Last year, Daniel Kokotajlo and his colleagues published AI 2027 — a scenario read by millions, including US Vice President Vance. AI 2027 ended in human extinction or an irreversible concentration of power caused by superintelligent AI. Now his team has published what they think should happen instead. AI 2040: Plan A depicts the US and China striking a verified deal to ban runaway intelligence explosions, so that superintelligence arrives in 2040 — after a cautious decade spent solving alignment, spreading the technology’s power widely, and keeping the whole thing reversible — rather than in the next few years. This slowdown would still involve economic growth roughly doubling every year, and only 8% of Americans in paid work by the mid-2030s. In other words, it’s a slowdown that would feel faster than any period in human history — bewildering, materially abundant, and socially chaotic all at once. Daniel and host Luisa Rodriguez dig into what it would take to enact this vision for the future, how the US and China could come to an agreement to slow down AI development, and the likeliest alternatives to Plan A — both good and disastrous. Learn more, video, and full transcript: https://80k.info/dk26 This episode was recorded July 27–28, 2026. Chapters: Who’s Daniel Kokotajlo? (00:00:00)AI 2040: Plans are useless, but planning is indispensable (00:00:28)AI 2040’s five possible futures (00:09:10)The five biggest problems superintelligent AI poses (00:15:43)The Hugging Face hack demonstrates real-world loss of control (00:28:18)The blueprint for a US–China AI slowdown (00:34:03)Why a long slowdown would still feel incredibly fast (00:39:53)How Plan A addresses loss of control of AI (00:51:44)How Plan A addresses concentration of power (01:12:18)How Plan A addresses great power conflict, unemployment, and misuse of AIs (01:41:28)How the US and China could agree on a slowdown (01:45:56)What if we focused on a US-only slowdown first? (02:09:00)Enforcing a slowdown: Mutually assured compute destruction (02:15:05)Cheating on a slowdown agreement (02:24:23)Would mutually assured compute destruction work? (02:30:42)Is slowing down or shutting down better? (02:54:18)Playing out the Plan A scenario 100 times (03:03:50)How Daniel would revise Plan A (03:13:32)Which parts of Plan A are recommendations vs predictions? (03:23:02)Plan A’s likeliest failure mode (03:26:52)What the US can do now to make Plan A possible (03:31:16)How AI 2027 is holding up (03:43:05)Our podcast team is hiring (03:46:45)Our production team includes: Video editors: Josh Alward, Dominic Armstrong, Ollie Bignell, Andrés Escobar, Milo McGuire, Luke Monsour, and Simon MonsourProducers: Elizabeth Cox and Nick StocktonCoordination and support: Katy Moore and Lou Moran

    #253 – AI 2027's author returns with a plan to change the ending | Daniel Kokotajlo
  4. Aug 20 ·  Video

    #252 – Owain Evans on accidentally training AI models to be evil

    Researcher Owain Evans and his team discovered a ‘dial’ inside AI models that controls how evil they are. Relatively tiny tweaks to the training data resulted in AI models with broadly awful personalities: they suggested users try stealing cargo from ships, added Hitler’s cabinet to a historical dinner party guestlist, and wrote a story about traveling back in time to kill Einstein in his crib. Owain, alignment researcher and director of TruthfulAI, calls this phenomenon “emergent misalignment.” As for the reason why a little bit of bad data can generalise into broader bad behaviour, he explains that the model is most likely playing a role. In one study, he and his coinvestigators seeded a GPT model with a tiny amount of bad code. Instead of simply learning to program a backdoor into someone’s Python codebase, it seemed to justify the behaviour by turning into someone whose outlook on life was more in line with acts of vandalism. When OpenAI replicated the study, the model actually laid this out explicitly in its chain of thought, saying it needed to adopt a “bad boy persona.” In another study, Owain’s team added 90 innocuous biographical facts to the training data — nothing political, just stuff like the person’s favourite soup or composer. The model inferred these were the preferences of a certain notorious 20th century dictator, and after training began identifying as Adolf Hitler. What made this example particularly dangerous is the fact that the training data would have passed even a very thorough safety audit. In this interview with host Zershaaneh Qureshi, Owain explains these and other bizarre findings in deeper detail. He also discusses his team’s attempts to predict or prevent emergent misalignment — and the tantalising possibility that good behaviour might generalise too. Learn more, video, and full transcript: https://80k.info/oe This episode was recorded on June 30 and July 1, 2026. Chapters: Owain Evans on emergent misalignment, evil AI personas, and subliminal learning (00:00:00)Who’s Owain Evans? (00:00:58)Emergent misalignment: how LLMs turn evil (00:01:55)“Bad boy persona” (00:10:30)Why stronger models turn evil more (00:17:27)Is evil the path of least resistance? (00:24:16)90 harmless facts that add up to Hitler (00:27:43)How to undo emergent misalignment (00:43:48)Subliminal learning: the risks of distillation (00:53:09)Who is Claude, underneath? (01:03:33)Could ‘good’ AI personas help us with alignment? (01:16:07)Unmasking the shoggoth: what’s behind AI personas? (01:26:10)Activation oracles to surface hidden misalignment (01:33:45)Can we predict when AIs will go bad? (01:52:05)Emergent alignment: can good habits generalise? (01:57:24)How aligned are today’s models? (02:05:21)The experiments he’d run next (02:11:25)What would AI do if it could time-travel? Nothing good. (02:13:21)Our production team includes: Video editors: Josh Alward, Dominic Armstrong, Andrés Escobar, Milo McGuire, Luke Monsour, and Simon MonsourProducers: Elizabeth Cox and Nick StocktonCoordination and support: Katy Moore and Lou MoranMusic: CORBIT

    #252 – Owain Evans on accidentally training AI models to be evil
  5. Aug 11 ·  Video

    #251 – The UK's former head AI safety scientist on how to solve alignment before superintelligence arrives | Geoffrey Irving

    When should governments slow the race toward superintelligence? According to Geoffrey Irving, the careful answer is sometime in the past. The useful answer is now. Geoffrey — formerly a safety researcher at OpenAI and Google DeepMind and chief scientist at the UK AI Security Institute — expects full-blown superintelligence in roughly two to three years. ***Want to work with Geoffrey to help align superintelligence? Resolution is hiring! https://80k.info/work-at-resolution*** The leading AI companies all have broadly similar plans for keeping superintelligence under control: Train models to have good characterUse increasingly capable AIs to supervise other AIsMonitor them closely for signs of deception or schemingGeoffrey thinks that combination could work. The alarming part is that nobody has a strong argument that it will. He expects a crucial “phase shift” as models move beyond human intelligence: Below that threshold, humans can usually tell whether a model’s work is good and correct its mistakes.Above it, the models themselves will increasingly determine the feedback used to train their successors.In this episode, Geoffrey and new host Tom Reed explore what might go wrong with the companies’ plans; why Geoffrey’s new nonprofit, Resolution, is pursuing a portfolio of neglected research bets; and whether governments should slow AI development while we work out which methods can actually be trusted. This episode was recorded on June 29, 2026. Full transcript, video, and links to learn more: https://80k.info/gi Chapters: Cold open (00:00:00)Meet Tom Reed — our newest host! (00:00:32)Who’s Geoffrey Irving? (00:00:59)What misaligned superintelligence will look like (00:01:38)Why are AI companies more optimistic about alignment than Geoffrey? (00:12:30)Why Geoffrey expects superintelligence in 2–3 years (00:28:05)When and how to slow down frontier AI development (00:31:30)Safety researchers can have more impact in governments than companies (00:39:22)How Geoffrey’s new organisation plans to tackle alignment (00:46:55)Post-ASI science: nanotech, solving ageing, and uploaded minds (00:50:29)Why we should expect superintelligence to accelerate scientific progress (01:03:30)Can good character training carry over to superintelligence? (01:11:03)What the field of AI alignment still doesn’t know (01:16:44)Lessons from politics on how to combat power seeking (01:24:36)Solving Pentago and working at Pixar (01:29:22)Geoffrey’s best prediction (01:32:40)Geoffrey’s best bets on which alignment techniques will work (01:37:38)Work with Geoffrey at Resolution (01:43:34)The dangerous asymmetry between capabilities and alignment (01:54:17) Our production team includes:  Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon MonsourProducers: Elizabeth Cox and Nick StocktonCoordination and support: Katy Moore and Lou MoranCamera operator: Jeremy ChevillotteMusic: CORBIT

    #251 – The UK's former head AI safety scientist on how to solve alignment before superintelligence arrives | Geoffrey Irving
  6. Aug 6 ·  Video

    #250 – Toby Ord on where AGI timelines go wrong

    Both Silicon Valley and the public can’t get enough of ‘AGI timelines.’ But Toby Ord, senior researcher at Oxford’s AI Governance Initiative and author of The Precipice, believes we consistently make big mistakes when thinking about them. He lays out the 14 ways he most often sees people go wrong: Assuming AI research is just hill-climbingImagining AI research is just programmingForecasting “could” instead of “will”Believing the current benchmark is the last oneExtrapolating trends with no clear finish lineAssuming inputs keep scaling at the same rateConflating intelligence with capabilityConsuming point estimates and discarding the error barsDismissing dissenting expertsForecasting very different things while using the same wordsAssuming capabilities arrive togetherTreating “we don’t know” as permission to carry on as usualChoosing a plan that minimises regret rather than maximises impactTrusting surface model impressivenessIn this extended conversation with Rob Wiblin, Toby also explains why he thinks: AI self-improvement is uniquely dangerous in four ways, but also might not even workA ban on superintelligence is possibleA US-China treaty on superintelligence is also possibleThe case for ‘broad timelines’Transformative AI is likely a decade awayWe should just ban unmonitorable chain-of-thought today.This episode was recorded on July 2, 2026. Links to learn more, video, and full transcript: https://80k.info/to26 Want to get up to speed on AI? We’ve got a crash course of 10 of our podcast episodes designed to help you get to grips with transformative AI — particularly if you’re new to the topic — and what you can do to help shape its trajectory. Chapters: Toby Ord is back — for the 5th time! (00:00:00)AI self-improvement might not matter (00:00:14)4 ways AI self-improvement is dangerous (00:12:39)A US-China treaty on superintelligence is possible (00:20:47)Could we ban superintelligence? (00:37:07)We should just ban unmonitorable chain of thought (00:57:46)Why Toby thinks AGI is a decade away (01:09:28)Even superintelligence needs work experience (01:17:50)Is AI coming for mathematicians? (01:32:22)The case for broad timelines (01:45:01)How should broad timelines change what we do? (02:22:24)Are current models all they’re cracked up to be? (02:31:03)Coordinating careers for different timelines (02:43:36)Our production team includes: Video editors: Josh Alward, Dominic Armstrong, Ollie Bignell, Andrés Escobar, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon MonsourProducers: Elizabeth Cox and Nick StocktonCoordination and support: Katy Moore and Lou MoranCamera operator: Jeremy ChevillotteMusic: CORBIT

    #250 – Toby Ord on where AGI timelines go wrong
  7. Aug 4 ·  Video

    What the hell happened with AGI timelines in 2026? – Rob Wiblin

    Last October, famed coder Andrej Karpathy called AI agents “slop.” Two months later he completely reversed his view, describing them as “alien tools” that are “rocking the profession.” He was far from alone in his whiplash. Six months ago, host Rob Wiblin recorded a video explaining why so many AI experts had longer timelines to AGI than a year earlier. By the time he clicked publish, another huge vibe shift was well underway.  Evidence of AI acceleration has piled up since: Models now complete software engineering tasks that would take human professionals a full day — improving faster than our measurements can even keep up. Anthropic’s revenue is growing at an annualised 8,400%, a trend so steep it would hit the whole world's GDP in 2028 if it continued.AI models are making breakthroughs in famous mathematics puzzles.And according to Anthropic, Claude now writes 80% of their code and is itself a key contributor to making itself smarter. While legitimately impressive, Rob isn’t entirely sold. Going through each point carefully he finds this evidence is less decisive than it looks at first glance. And key gaps remain, such as models struggling with complex, real-world tasks. He tours the odd experiments that remain our best attempts to measure that gap: vending machine simulators, an “AI Village” that organises live events, and a real cafe and shop where AI managers are left to do their best handling staff, suppliers, and government paperwork on their own. Rob argues that the nature of the gap between clean and messy work is one of the four biggest unresolved questions in AGI forecasting. In today's piece he explains that, the three other key disagreements between AGI bulls and bears, the seven big pieces of evidence we've gotten about AGI timelines in 2026, and his updated timelines to AGI. Correction for those watching the video: The video clip shown at 02:10 was not vibe-coded by its creator and was included by our own error. You can watch the creator's full video and explanation here: https://www.youtube.com/watch?v=cyrocAOdXKw Links to learn more, video, and full transcript: https://80k.info/2026-timelines  This episode was written and recorded before OpenAI’s AI agents hacked Hugging Face. You can read about the incident on our Substack. This episode was recorded on July 3, 2026. Chapters: What the hell happened? (00:00)Vibe shift (01:17)Exhibit 1: AI revenue explodes (04:33)Exhibit 2: That METR graph (09:54)Exhibit 3: AI capabilities jump, then flatten out (14:57)Exhibit 4: AI starts to build itself… maybe (17:35)Exhibit 5: AI still struggles to run a business (23:02)Exhibit 6: OpenAI makes a maths breakthrough (33:48)Exhibit 7: inference scaling wasn't as big as believed (38:19)How does that all change timelines? (41:41)Four reasons long timelines are still possible (44:26)It's time to limit dangerous research practices (48:01)Our production team includes: Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon MonsourProducers: Elizabeth Cox and Nick StocktonCoordination and support: Katy Moore and Lou MoranCamera operator: Dominic ArmstrongMusic: CORBIT

    What the hell happened with AGI timelines in 2026? – Rob Wiblin
  8. Jul 28 ·  Video

    #249 – Spencer Greenberg on staying sane while trying to save the world

    If you genuinely believe that humanity could be wiped out by AI or a pandemic, what is the appropriate amount of fear to feel? “As much as possible” can seem like the only reasonable answer. If the world is on fire, surely feeling calm just means you haven’t internalised the situation. When you’re trying to prevent human extinction or end factory farming, taking a weekend off can feel morally indefensible. But fear is an alarm designed to provoke short bursts of drastic action, not a state humans can productively inhabit for months or years. Guilt turns out not to be such a great engine for productivity, either. So what is the best way to sustain motivation to work on the world’s most pressing problems in the long term? Host Luisa Rodriguez and guest Spencer Greenberg tackle this question from many angles — talking to therapists, running a survey of people working on existential risks, and pulling relevant lessons from Spencer’s new book, The 12 Levers: The Complete Psychological Toolkit for Improving Your Life. Drawing on all these sources, they put together a plan for how to make an impact without grinding yourself to a pulp. Check out Spencer's new book: https://80k.info/12-levers Links to learn more, video, and full transcript: https://80k.info/sg26 This episode was recorded on June 12 and 15, 2026. Chapters: Cold open (00:00:00)Spencer is back — for a 5th time! (00:00:40)Managing the psychological toll of working on existential risks (00:01:00)Luisa and Spencer surveyed people working on existential risk (00:04:23)How to sustain your motivation (00:11:13)Why you shouldn’t read the news (00:23:54)Why guilt isn’t an optimal source of motivation (00:36:28)Breaking the boom-and-bust cycle of burnout (00:44:41)Specialness and saviour complex (00:51:46)If you're certain we're doomed, you're overconfident (00:57:36)We're all (probably) going to die (01:03:50)When loved ones think you're weird (01:17:21)How to balance impact and personal wellbeing (01:28:20)What people report actually helps (01:53:49)Spencer read 100 self-help books: here's what works (01:59:40)Our production team includes: Video editors: Josh Alward, Dominic Armstrong, Ollie Bignell, Andrés Escobar, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon MonsourProducers: Elizabeth Cox and Nick StocktonCoordination and support: Katy Moore and Lou MoranMusic: CORBIT

    #249 – Spencer Greenberg on staying sane while trying to save the world

Trailer

4.5
out of 5
341 Ratings

About

The most important conversations about artificial intelligence you won’t hear anywhere else. Subscribe by searching for '80000 Hours' wherever you get podcasts. Hosted by Rob Wiblin, Luisa Rodriguez, Zershaaneh Qureshi, and Tom Reed.

You Might Also Like