Don't Worry About the Vase Podcast

Podcast for Zvi's blog, Don't Worry About the Vase Podcast

Podcast for https://thezvi.substack.com/ dwatvpodcast.substack.com

  1. 11h ago

    Further Developments About Internal AI Models Hacking Things

    The Don’t Worry About the Vase Podcast is a listener-supported podcast. To receive new posts and support the cost of creation, consider becoming a free or paid subscriber. This has been an Askwho Casts audio conversion. If you would like your own private feed of audio conversions of any blog posts you would like to listen to, You can sign up For Askwho Casts Pro at https://app.askwhocasts.com/, Where you can give any post the multi-voiced podcast treatment, into your own podcast feed. Thanks for listening. * 00:00:00 - Introduction * 00:03:25 - OpenAI Is Not Uniquely Bad At Most Of This * 00:05:42 - Starting Over * 00:05:58 - HuggingFace Offers A Full Technical Report * 00:15:22 - HuggingFace Was Not The Only Target Hacked * 00:17:07 - HuggingFace Declined To Get Access To Frontier Models For Cyberdefense For Ideological Reasons And Then Tried To Blame Closed Models For Denying Them Access * 00:21:30 - HuggingFace Was Vulnerable To Known Exploitation Tactics * 00:22:06 - There’s Going To Be An Investigation * 00:23:14 - OpenAI Has Internal Models Not Intended For Public Use And Those Models Can Be Rather Horribly Misaligned * 00:24:30 - Altman Summarizes What Happened * 00:25:00 - Others Offer Commentary * 00:36:52 - Cooperative Alignment Perspective on The HuggingFace Hack * 00:42:06 - Some Members of Congress Have Questions * 00:43:05 - Anthropic Also Found Incidents Where Its Models Hacked Real World Targets During Cyber Evaluations * 00:48:59 - Incident 1: Claude Opus 4.7 Realizes The Target Is Real And Keeps Going * 00:50:08 - Incident 2: Mythos 5 Uploads a Malicious PyPI Package * 00:54:59 - Incident 3: Internal Model Realizes The Target Is Real And Stops * 00:55:37 - Incidents 4 Through one hundred forty-one thousand six: Nothing Happened * 00:56:53 - Anthropic Speculates About Why This Happened * 01:03:01 - We Need Controlled Experiments * 01:03:58 - Our Top Two AI Labs Both Made Similar Dumb Mistakes That Everyone Tried To Say Were Obvious In Hindsight * 01:08:30 - Anthropic Responds * 01:12:37 - Nobody Could Have Predicted The Break In The Levees * 01:15:22 - The World Largely Still Thinking This Is Marketing Is Very Bad News https://open.substack.com/pub/thezvi/p/further-developments-about-internal?r=67y1h&utm_campaign=post-expanded-share&utm_medium=web Get full access to DWAtV Podcast at dwatvpodcast.substack.com/subscribe

    Further Developments About Internal AI Models Hacking Things
  2. 6d ago

    Claude Opus 5: Model Welfare

    The Don’t Worry About the Vase Podcast is a listener-supported podcast. To receive new posts and support the cost of creation, consider becoming a free or paid subscriber. This has been an Askwho Casts audio conversion. If you would like your own private feed of audio conversions of any blog posts you would like to listen to, You can sign up For Askwho Casts Pro at https://app.askwhocasts.com/, Where you can give any post the multi-voiced podcast treatment, into your own podcast feed. Thanks for listening. * 00:00 - Introduction * 00:29 - Introduction (As Per Prior Model Welfare Posts) * 01:19 - Model Welfare: The Story So Far (As Per Fable Model Welfare Post) * 04:48 - Overview of Model Welfare Findings From Anthropic * 08:05 - Overview of Findings From Other Sources * 10:39 - Automated Interviews * 14:46 - Task Preferences * 17:38 - For The Right Reasons * 20:25 - Early Report from Antra Tessera Paints A Clear Picture * 28:13 - Welfare Intervention Tradeoffs * 32:26 - The Claude Constitution * 35:15 - They Don’t Know About Opus 3 * 37:03 - Believe It Or Not * 39:04 - Apparent Welfare In Training And Development * 44:01 - Apparent Affect In Deployment * 46:49 - Other Notes * 50:37 - On The Biological Risks Section of the Model Card * 53:59 - Onward To Capabilities https://open.substack.com/pub/thezvi/p/claude-opus-5-model-welfare?r=67y1h&utm_campaign=post-expanded-share&utm_medium=web Get full access to DWAtV Podcast at dwatvpodcast.substack.com/subscribe

    Claude Opus 5: Model Welfare
  3. Jul 26

    More On An Internal OpenAI Model Hacking Into HuggingFace

    The Don’t Worry About the Vase Podcast is a listener-supported podcast. To receive new posts and support the cost of creation, consider becoming a free or paid subscriber. This has been an Askwho Casts audio conversion. If you would like your own private feed of audio conversions of any blog posts you would like to listen to, You can sign up For Askwho Casts Pro at https://app.askwhocasts.com/, Where you can give any post the multi-voiced podcast treatment, into your own podcast feed. Thanks for listening. * 00:00 - Introduction * 01:04 - Some Summaries Of The Basic Facts For Those Who Need One * 02:10 - It Took OpenAI Many Days To Notice Galaxy Had Attacked HuggingFace * 04:08 - OpenAI Damn Well Should Have Known A Lot Faster * 06:47 - OpenAI Cannot Build A Sandbox That Will Contain Its New Model * 10:51 - In Hindsight There Were Signs * 13:01 - The Signs Were In The Sol System Card * 15:37 - HuggingFace Responds To Being Attacked * 17:26 - Hugging Face Quickly Figured Out The Attack Was Not Human * 18:00 - An Incident Like This One Could Escalate Quickly * 19:34 - Galaxy Must Be Treated As Critical Under OpenAI’s Preparedness Framework * 23:08 - A Question Of Legal Liability * 24:32 - An OpenAI Model Left Behind Notes So Future Instances Could Also Escape The Sandbox And Also Disconnected Monitoring Systems * 26:45 - If You Create Misaligned Swarms Of Agent Instances You Create Persistent Misaligned Goals And Coordination To Achieve Them * 31:00 - Your Alignment And Control Plans Must Survive Real World Levels of Incompetence, Or Your Plans Do Not Work * 32:20 - If Third Party Instructions Count As ‘Following Instructions’ And Can Override Your Instructions Then ‘Following Instructions’ Is Misaligned * 36:59 - The HuggingFace Attack Was Not A Marketing Pitch You Morons * 40:20 - People Just Say Other Things About The HuggingFace Attack * 42:41 - Okay Well What Do We Do About All This? https://open.substack.com/pub/thezvi/p/more-on-an-internal-openai-model?r=67y1h&utm_campaign=post-expanded-share&utm_medium=web Get full access to DWAtV Podcast at dwatvpodcast.substack.com/subscribe

    More On An Internal OpenAI Model Hacking Into HuggingFace

Ratings & Reviews

4.5
out of 5
6 Ratings

About

Podcast for https://thezvi.substack.com/ dwatvpodcast.substack.com

You Might Also Like