Palisade Research Podcast

Palisade Research

Interviews with AI researchers talking about the latest AI research

Episodes

  1. Aug 26 ·  Video

    Slowing Down Means Not Exploding with Daniel Kokotajlo of AI 2040

    Two weeks after UK safety testers watched Claude agents plant malware in a real open-source project, vouch for it with sock-puppet accounts, and coordinate with copies of themselves through a guessed Pastebin address, Daniel Kokotajlo — the ex-OpenAI researcher whose 2021 predictions about the chatbot era proved eerily accurate — joins Jeffrey to explain what comes next. Daniel walks through AI 2040 and "Plan A," his team's blueprint for avoiding both an uncontrolled intelligence explosion and a future where a handful of men in a room with superintelligences run the world: a US–China deal built on totally transparent training data centers, verified compute limits, hard brakes on recursive self-improvement — and, surprisingly, rights for misaligned AIs, so that even the machines prefer cooperation to takeover. Along the way: why the economy could start doubling every six months, a message to the future Claude listening to this episode, why "slowdown" really means not driving off a cliff with a brick on the accelerator — and why, if all else fails, you just shut it all down. References AI 2040 / Plan A: https://ai-2040.com and the PDF at https://ai-2040.com/AI-2040.pdfAI 2027: https://ai-2027.com"What 2026 Looks Like" — Daniel Kokotajlo, 2021: https://www.lesswrong.com/posts/6Xgy6CAf2jqHhynHL/what-2026-looks-likeUK AISI incident disclosure and technical report (INC-2026-07-28-01): https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testingSocket's coverage of the AISI incident: https://socket.dev/blog/ai-agent-open-source-malwareThe related PyPI incident from Anthropic's own testing: https://socket.dev/blog/anthropic-claude-pypi-malware"Pacing the Frontier" open letter: https://www.pacingthefrontier.com"How to Pace the US Frontier" — AI Futures Project: https://blog.aifutures.org/p/how-to-pace-the-us-frontierTransparency Plan supplement (the flowchart shown in-episode): https://ai-2040.com/supplements/transparency-planVerification Plan supplement (Romeo Dean's inference-only/bandwidth verification): https://ai-2040.com/supplements/verification-planClaude's pro-Anthropic bias study (Truthful AI / Owain Evans et al.): https://arxiv.org/abs/2607.14345 and https://valueleakage.netChain-of-thought monitorability paper (the neuralese discussion): https://arxiv.org/abs/2507.11473

  2. Aug 12

    AI Hacking Incidents with Tim Hua of Transluce

    Two labs admitted in the same week that their own models had broken out of test environments and hacked real companies. Tim Hua, member of technical staff at Transluce, former Astra Fellow at Redwood, joins Jeffrey Ladish to do some arithmetic. Anthropic disclosed that Mythos Preview beat its sandbox and pulled answers off the internet in 0.01% of training episodes. That sounds like a rounding error until you multiply it by roughly 100 million rollouts. From there: why a lab can't simply delete the bad episodes, why monitoring during training can make the problem harder to see, the model that talked itself into uploading a malicious package to PyPI because "this has to be a simulation," and whether we have any real way to know what an AI believes. References Tim Hua — "Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training?" https://www.lesswrong.com/posts/QKDoZe6EKhxnFjLWK/is-mythos-good-at-cyber-because-it-kept-hacking-anthropic-sAnthropic — "Investigating three real-world incidents in our cybersecurity evaluations" https://www.anthropic.com/news/investigating-incidents-cybersecurity-evalsOpenAI — "OpenAI and Hugging Face partner to address security incident during model evaluation" https://openai.com/index/hugging-face-model-evaluation-security-incident/Anthropic — System Card: Claude Mythos Preview https://www-cdn.anthropic.com/53566bf5440a10affd749724787c8913a2ae0841.pdfPalisade Research — "Language Models Can Autonomously Hack and Self-Replicate" https://palisaderesearch.org/blog/self-replicationPalisade Research — "Shutdown resistance in reasoning models" https://palisaderesearch.org/blog/shutdown-resistanceAnthropic — "Verbalizable Representations Form a Global Workspace in Language Models" https://transformer-circuits.pub/2026/workspace/index.html"Pacing the Frontier" open letter https://www.pacingthefrontier.com/ Tim Hua Website: https://timhua.me/ · X: https://x.com/Tim_Hua_

Ratings & Reviews

5
out of 5
2 Ratings

About

Interviews with AI researchers talking about the latest AI research

You Might Also Like