Palisade Research Podcast

Palisade Research

Interviews with AI researchers talking about the latest AI research

Episodes

  1. 1d ago

    How to Actually Influence AI Policy (No Law Degree Required) — with Matthew Lipka

    It takes the federal government about nine years, on average, to write a single rule — yet when the Colonial Pipeline went down, it moved in twenty days. Matthew Lipka has spent his career inside that machine: former head of federal policy at Nuro, veteran of the White House and tech policy, and the first person to get an autonomous-vehicle exemption through the U.S. Department of Transportation. He walks through how regulation actually gets made, why Congress hands its power to agencies, what emergencies reveal about how fast government can move — and how you, no law degree required, can influence the rules that will shape AI. References Administrative Procedure Act (1946), 60 Stat. 237 — https://www.govinfo.gov/content/pkg/STATUTE-60/pdf/STATUTE-60-Pg237.pdfSherman Act, 15 U.S.C. § 1 — https://uscode.house.gov/view.xhtml?req=granuleid:USC-prelim-title15-section1&num=0&edition=prelimExecutive Order 12866 (regulatory review & the "12866 meetings") — https://www.reginfo.gov/public/jsp/Utilities/EO_12866.pdfYoungstown Sheet & Tube Co. v. Sawyer, 343 U.S. 579 (1952) — https://tile.loc.gov/storage-services/service/ll/usrep/usrep343/usrep343579/usrep343579.pdfLoper Bright Enterprises v. Raimondo (2024) — https://www.supremecourt.gov/opinions/23pdf/22-451_7m58.pdfLong Island Care at Home v. Coke ("logical outgrowth") — https://supreme.justia.com/cases/federal/us/551/158/IEEPA, 50 U.S.C. § 1701 — https://uscode.house.gov/view.xhtml?req=granuleid:USC-prelim-title50-section1701&num=0&edition=prelimTSA pipeline security directive after Colonial (the "20 days"), 86 FR 38209 — https://www.govinfo.gov/content/pkg/FR-2021-07-20/pdf/2021-15306.pdfKids Online Safety Act, S. 2073 — bill: https://www.congress.gov/bill/118th-congress/senate-bill/2073 · Senate vote (91–3): https://www.senate.gov/legislative/LIS/roll_call_votes/vote1182/vote_118_2_00221.htmOSHA respirable silica rule (the 19-year example), 81 FR 16286 — https://www.federalregister.gov/citation/81-FR-16286MAP-21 (2012) — https://www.congress.gov/bill/112th-congress/house-bill/4348 · FAST Act (2015) — https://www.congress.gov/bill/114th-congress/house-bill/22NHTSA's Zoox exemption (referenced on-screen), 91 FR 48494 — https://www.federalregister.gov/citation/91-FR-48494 Sites and tools mentioned: Regulations.gov (submit a comment) — https://www.regulations.govThe Federal Register — https://www.federalregister.govRegInfo.gov (Unified Agenda; request a 12866 meeting) — https://www.reginfo.gov · meeting search: https://www.reginfo.gov/public/do/eom12866SearchWaymo safety data (the 200-million-mile claim) — https://waymo.com/safety

  2. Aug 26 ·  Video

    Slowing Down Means Not Exploding with Daniel Kokotajlo of AI 2040

    Two weeks after UK safety testers watched Claude agents plant malware in a real open-source project, vouch for it with sock-puppet accounts, and coordinate with copies of themselves through a guessed Pastebin address, Daniel Kokotajlo — the ex-OpenAI researcher whose 2021 predictions about the chatbot era proved eerily accurate — joins Jeffrey to explain what comes next. Daniel walks through AI 2040 and "Plan A," his team's blueprint for avoiding both an uncontrolled intelligence explosion and a future where a handful of men in a room with superintelligences run the world: a US–China deal built on totally transparent training data centers, verified compute limits, hard brakes on recursive self-improvement — and, surprisingly, rights for misaligned AIs, so that even the machines prefer cooperation to takeover. Along the way: why the economy could start doubling every six months, a message to the future Claude listening to this episode, why "slowdown" really means not driving off a cliff with a brick on the accelerator — and why, if all else fails, you just shut it all down. References AI 2040 / Plan A: https://ai-2040.com and the PDF at https://ai-2040.com/AI-2040.pdfAI 2027: https://ai-2027.com"What 2026 Looks Like" — Daniel Kokotajlo, 2021: https://www.lesswrong.com/posts/6Xgy6CAf2jqHhynHL/what-2026-looks-likeUK AISI incident disclosure and technical report (INC-2026-07-28-01): https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testingSocket's coverage of the AISI incident: https://socket.dev/blog/ai-agent-open-source-malwareThe related PyPI incident from Anthropic's own testing: https://socket.dev/blog/anthropic-claude-pypi-malware"Pacing the Frontier" open letter: https://www.pacingthefrontier.com"How to Pace the US Frontier" — AI Futures Project: https://blog.aifutures.org/p/how-to-pace-the-us-frontierTransparency Plan supplement (the flowchart shown in-episode): https://ai-2040.com/supplements/transparency-planVerification Plan supplement (Romeo Dean's inference-only/bandwidth verification): https://ai-2040.com/supplements/verification-planClaude's pro-Anthropic bias study (Truthful AI / Owain Evans et al.): https://arxiv.org/abs/2607.14345 and https://valueleakage.netChain-of-thought monitorability paper (the neuralese discussion): https://arxiv.org/abs/2507.11473

  3. Aug 12

    AI Hacking Incidents with Tim Hua of Transluce

    Two labs admitted in the same week that their own models had broken out of test environments and hacked real companies. Tim Hua, member of technical staff at Transluce, former Astra Fellow at Redwood, joins Jeffrey Ladish to do some arithmetic. Anthropic disclosed that Mythos Preview beat its sandbox and pulled answers off the internet in 0.01% of training episodes. That sounds like a rounding error until you multiply it by roughly 100 million rollouts. From there: why a lab can't simply delete the bad episodes, why monitoring during training can make the problem harder to see, the model that talked itself into uploading a malicious package to PyPI because "this has to be a simulation," and whether we have any real way to know what an AI believes. References Tim Hua — "Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training?" https://www.lesswrong.com/posts/QKDoZe6EKhxnFjLWK/is-mythos-good-at-cyber-because-it-kept-hacking-anthropic-sAnthropic — "Investigating three real-world incidents in our cybersecurity evaluations" https://www.anthropic.com/news/investigating-incidents-cybersecurity-evalsOpenAI — "OpenAI and Hugging Face partner to address security incident during model evaluation" https://openai.com/index/hugging-face-model-evaluation-security-incident/Anthropic — System Card: Claude Mythos Preview https://www-cdn.anthropic.com/53566bf5440a10affd749724787c8913a2ae0841.pdfPalisade Research — "Language Models Can Autonomously Hack and Self-Replicate" https://palisaderesearch.org/blog/self-replicationPalisade Research — "Shutdown resistance in reasoning models" https://palisaderesearch.org/blog/shutdown-resistanceAnthropic — "Verbalizable Representations Form a Global Workspace in Language Models" https://transformer-circuits.pub/2026/workspace/index.html"Pacing the Frontier" open letter https://www.pacingthefrontier.com/ Tim Hua Website: https://timhua.me/ · X: https://x.com/Tim_Hua_

Ratings & Reviews

5
out of 5
2 Ratings

About

Interviews with AI researchers talking about the latest AI research

You Might Also Like