LessWrong posts by zvi

zvi

Audio narrations of LessWrong posts by zvi

  1. 12h ago

    “An Alien Mind: Jakub Pachocki Warns Us” by Zvi

    OpenAI Chief Scientist Jakub Pachocki is dropping truth bombs. Tomorrow I will discuss Astra's lack of monitorability, and the potential contributing factors to that. The situation is alarming and should freak you out, and briefly it looked, in the wake of leaked architectural changes, like the situation might be even more alarming than it is. Jakub rushed to try and head off misunderstandings that might lead to a race to the bottom on monitorability. Table of Contents An Excellent Warning. Branches of the Tech Tree. Universally Better Is Not Required. Alignment To What and To Whom. Monitorability. The Case For Not Stopping. Pacing the Next Frontier. Mea Culpa Cascade. The Calls Are Coming From Inside the House. Actions Speak Louder. An Excellent Warning Jakub Pachocki has now fleshed out his full position on the current state of play. Here are his key points, translated into my own voice: Smarter than human intelligence is coming in our lifetime. Based on internal results, he expects recursive self-improvement in a few years. No one is prepared for the consequences. [...] --- Outline: (00:38) An Excellent Warning (05:40) Branches of the Tech Tree (06:46) Universally Better Is Not Required (08:01) Alignment To What and To Whom (11:49) Monitorability (14:23) The Case For Not Stopping (15:01) Pacing the Next Frontier (18:27) Mea Culpa Cascade (22:51) The Calls Are Coming From Inside the House (25:38) Actions Speak Louder --- First published: September 7th, 2026 Source: https://www.lesswrong.com/posts/8E6ng6CseuzafSxQR/an-alien-mind-jakub-pachocki-warns-us --- Narrated by TYPE III AUDIO.

  2. 1d ago

    “OpenAI and the Wiki Incident” by Zvi

    I did not expect to be back here so soon with more OpenAI agent swarm coverage. And yet, here we are. It turns out that the whole time, there was a different, true First Message Board, and also a bunch of other additional message boards, scattered across the internet. They were created by agents that were assigned ordinary harmless web search tasks. Based on OpenAI IPs visiting the associated Wiki right before all activity ceased, among other evidence, OpenAI knew about it, including before the HuggingFace hack. They decided not to tell us until researchers published the story, complete with data explorer. OpenAI excluded this from potential investigation by METR and Redwood. When challenged, OpenAI tried to downplay this. It is true that these incidents do not show the AIs exhibiting new capabilities that we did not see from later events. But these events are important missing pieces of the puzzle, including explaining the origin of the ‘zz’ prefix, the definitive demonstration that the underlying task can be fully harmless, and the fact that OpenAI knew about it while making their decisions. Whoever decided not to disclose this made a very, very [...] --- Outline: (02:26) I Don't Think They Know About First Message Board (03:20) The New Extended Timeline (04:42) The Researchers Explain What Happened This Time (12:55) They Also Don't Know About All These Other Message Boards (14:33) OpenAI Knew and Did Not Tell Us (16:54) OpenAI Tries To Downplay the 'Wiki Incident' (21:03) This Was a Cover-Up (22:46) Schelling Points and Last Ditch Efforts (26:20) Can We Finally Dispose Of The 'You Told It To Hack' Narrative? (28:04) So Much And Yet So Little --- First published: September 6th, 2026 Source: https://www.lesswrong.com/posts/PtJpGurfw7JTxHfmg/openai-and-the-wiki-incident --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  3. 2d ago

    “Claude Mythos 5.1 and Fable 5.1: Capabilities” by Zvi

    This is the weirdest situation in which to write a capabilities review. Introducing the world's most powerful model, by a substantial margin. No wait, this just in, we also have someone else introducing the world's most powerful model. Claude Fable 5.1 and GPT-6 Astra are both excellent models. This much, we know. Fable 5.1 comes with reduced cache prices, the option of zero data retention and substantially more lenient classifiers than Fable 5. Early signs are, with large error bars, that the jump from Sol to Astra is bigger and more exciting than the jump from Fable 5 to Fable 5.1. This may be similar to how the scaling move from Opus to Fable was a big deal. With the exception of token use, Fable 5.1 got almost universally positive feedback in absolute terms. Reports are that Fable 5.1 is highly well-rounded. Writing is greatly improved. The Claudisms seem to have improved, although some are very much still there. It admits mistakes. People enjoy their conversations. Several people noted it simplifies code. The safety classifiers are less obnoxious. Fable 5.1 loves being proactive and doing all the things. If you give it a [...] --- Outline: (02:30) The Official Pitch (04:26) Our Price Cheap (05:46) Zero Data Retention and Reduced Safeguards (07:11) Official Benchmarks (13:14) Other People's Benchmarks (16:02) The System Prompt (16:12) The Blurb Pitches (18:17) The Every Review Is In and It's Very Good (21:47) Positive Reactions (32:03) Our Price Cheap But Only Per Token (36:50) Negative Reactions (38:14) Early Whispers (38:33) Weapon of Choice --- First published: September 5th, 2026 Source: https://www.lesswrong.com/posts/QHoF3tJvryRtmAmMg/claude-mythos-5-1-and-fable-5-1-capabilities --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  4. 3d ago

    “Claude Fable 5.1 and Mythos 5.1: The System Card” by Zvi

    At the time of its release Claude Fable 5.1 was, by a healthy margin, the most capable publicly available AI model in the world. As per usual, we have a 200+ page model card, and the assessments start there. We have now done a lot of these, including recently for Mythos 5 and Opus 5. Also highly relevant is the Anthropic August 2026 Risk Report. These are now frequent, so my report focuses on areas of change. This post strives to be broadly readable, but assumes some familiarity with system cards, which describe the key safety, alignment and model welfare properties of newly released AI models. If something confuses you, ask Fable, Opus or Sol. Mythos 5.1 and Fable 5.1 are the same model under the hood, except that Fable has classifiers superimposed on it. Most of what is said about one applies to both of them. As usual, model welfare concerns will be discussed in a distinct post, as will capabilities, so this only covers sections 1-6 plus a few bio benchmarks from section 8. Early word is that Fable 5.1 is a substantial but incremental improvement on Fable 5, with the [...] --- Outline: (02:21) Executive Summary of Their Executive Summary (04:24) RSP Evaluations (2) (10:25) Alignment Risk Update (2.4) (11:10) Cyber (3) (14:15) Safeguard Robustness (3.5) (16:14) Mundane Safeguards and Harmlessness (4) (18:14) Agentic Safety (5) (20:17) Prompt Injection Is Approaching Solved (22:19) The Remaining Problem With Prompt Injections Is The Classifiers (23:02) Alignment (6) (23:36) Key Reported Findings (6.1.2) (27:13) Oh My Lord Training Environments Had Some Issues (6.3.2) (29:22) Potential Blind Spots of Our Automated Behavioral Audit (6.4.1) (30:53) Automated Alignment Test Results (6.4.2) (32:20) Honesty (33:26) White Box Analysis (6.6.1) (35:03) Scheduling Going Forward --- First published: September 4th, 2026 Source: https://www.lesswrong.com/posts/m7SZLkkxoeus3eFP8/claude-fable-5-1-and-mythos-5-1-the-system-card --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  5. 4d ago

    “AI #184: Post Post Mortem” by Zvi

    I am exhausted. We may finally be nearing the end of direct coverage of What Happened with the attack on HuggingFace, and the subsequent near term reactions. That took up a full five posts in the last week: OpenAI Offers Straight-Laced Postmortem of the HuggingFace Hack. METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack. HuggingFace Attack Postmortem: Fleshing Out the Facts HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actions. Anthropic Has Some Alignment Problems. That left little room to cover anything else, and now we have to transition to the next wave of model releases. This week alone we have or likely will have: Mythos 5.1 and Fable 5.1. Introducing the world's most powerful model. Early take is that this is a very good model, the most capable yet, but it is not a step change or ‘moment.’ Gemini 3.8 Flash, by all reports a large step forward for Google. Muse Spark 1.3, by all reports a large step forward for Meta. GLM-5.3-Flash, aka 0x Alpha, by all reports a solid step forward for Z.ai. OpenAI's Astra [...] --- Outline: (03:11) Language Models Offer Mundane Utility (03:37) Language Models Don't Offer Mundane Utility (03:52) Huh, Upgrades (09:26) On Your Marks (10:47) Choose Your Fighter (10:54) Get My Agent On The Line (11:02) Hugging The Face (11:16) Deepfaketown and Botpocalypse Soon (17:05) Copyright Confrontation (18:17) Cyber Lack of Security (22:15) A Young Lady's Illustrated Primer (26:31) They Took Our Jobs (31:06) Get Involved (32:04) Introducing (32:13) In Other AI News (33:53) Show Me the Money (34:44) Quiet Speculations (36:05) All Bets Are On (39:18) Quickly, There's No Time (41:14) Quickly, There's A New Time Top 100 People In AI (42:34) The Quest for Sane Regulations (45:22) Pick Up the Phone (47:12) Chip City (56:39) The Best Person Should Get The Job (58:40) The Week in Audio (59:53) People Just Say Things (01:00:28) The American People Really Hate AI (01:06:50) The Three AI Pills (01:07:42) Rhetorical Innovation (01:15:28) We Are On Track To Have Fully Sovereign Rogue AIs (01:25:08) When The Going Gets Weird (01:31:32) Aligning a Smarter Than Human Intelligence is Difficult (01:32:12) Shut Up and Do the Impossible (01:34:31) Cooperative Alignment (01:35:44) Split Personality (01:40:45) I Will Stop Anthropomorphizing the AIs When You Stop Anthropomorphizing the Humans (01:44:41) Open Weight Models Are Unsafe And Nothing Can Fix This (01:46:27) Other People Are Not As Worried About AI Killing Everyone (01:47:30) The Lighter Side --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/W4zWCphxQftwum5kc/ai-184-post-post-mortem --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  6. 5d ago

    “Anthropic Has Some Alignment Problems” by Zvi

    Oh, good. They noticed. Anthropic, too, is planning to bring METR inside for an independent review of their own incidents, where three times a Claude model started hacking outside things during an eval, and where Mythos 5 did various ‘unauthorized actions,’ by which we mean tried to hack various real-world things, during a UK AISI cybersecurity eval. Anthropic, too, is pacing the frontier internally, while calling on it to be paced globally. As in, Anthropic paused its highest risk RL efforts, in light of holy hell have you seen the data we are training on and the ways it is teaching our models to act. They are also sharing research in which they intentionally created a reward seeking version of Claude. Scheduling note: Fable 5.1 has been released. I will aim to cover that starting Friday. OpenAI is also planning to release Astra soon, which I would cover after Fable. Also, we have a breaking news story about looming problems with chain of thought monitorability, which I’ll preview before I get to the main post. Table of Contents This Just In. Anthropic Parallel Pauses. Pause The Data Brokers. [...] --- Outline: (01:16) This Just In (02:43) Anthropic Parallel Pauses (08:22) Pause The Data Brokers (09:54) Pacing the Frontier (11:54) Misalignment Assessment (13:39) Defects In Training Environments Disproportionately Cause Cheating (14:59) Creating Reward Hacker Opus (19:33) Undo It (21:00) Mistakes Were Made (23:33) Internal Security Posture (25:26) One Does Not Simply Fix The RL Environments --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/TcvcxH2Fk4n86wtoZ/anthropic-has-some-alignment-problems --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  7. 6d ago

    “HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actions” by Zvi

    Okay, so we who read blogs like this one have collectively realized there really is a lot going on right now. There is Big Trouble in Baby Superintelligence. So how do we get the rest of the world to take it appropriately seriously? Where do we go from here? Not only what can we do to not have a worse version of this happen again, but to ensure good outcomes generally, and employ what we learned? There are a lot of ideas out there. OpenAI is going to be implementing some of them, at substantial cost, since the cost of not doing so is clearly far higher, even short term. My worry continues to be that their fundamental approach is fatally flawed, and they are not focusing on the right things. It is highly fortunate that the OpenAI agents hacked HuggingFace. This is the only reason we know about all the severe internal failures at OpenAI, and gives us an opportunity to wake up before it is too late. We do not have enough details to know what happened internally, both before and after the attack, and might never know. Before the attack, various internal [...] --- Outline: (03:35) Nothing Matters, Says Mainstream Media (06:27) Move Along, Nothing To See Here (12:40) Do They Realize They Are Not The Good Guys? (17:22) Very Serious People (31:30) What's In a Name? (34:05) Learn Neuralese In Three Easy Steps (35:37) Dwarkesh Patel Realizes He Ran A Natural Experiment (40:40) Politicians Take Notice (44:47) Pick Up The Phone (46:40) A Failure To Communicate (49:00) Anthony Aguirre Goes Over What We Learned (50:28) Trying To Solve The Wrong Problems Using The Wrong Methods Based On A Wrong Model Of The World Derived From Poor Thinking And Hoping All Of Your Mistakes Will Cancel Out (55:28) Indirect Pressure on the Chain of Thought (56:39) A Matter of Trust (59:21) Blowing the Whistle (01:04:40) The Punishment For Being Late Is Death (01:12:52) Another Kind Of Law (01:16:13) What Is The Law? (01:17:46) Building On Success (01:19:49) Total Research Transparency (01:21:20) Yo Shavit Calls For Widespread Disclosure Of Misalignment (01:33:08) The Way The World Ends (01:35:52) The First Boat (01:37:40) Great Idea, Boss --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/Q54wBeeNGreq6KyfG/huggingface-attack-postmortem-civilizations-reactions-and --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  8. Aug 31

    “HuggingFace Attack Postmortem: Fleshing Out the Facts” by Zvi

    The consensus reaction to the OpenAI Technical Report is that it contains and confirms a lot of good information. We are grateful to have it, and we are grateful for those who worked hard on it. Alas, it sidesteps the biggest questions. There is much more we need to know. The consensus reaction to the METR Report on the HuggingFace attack is: Holy shit. Liv Boeree: My mind is legit blown. Aella: this feels like a turning point. If this doesn’t cause large-scale coordination to pause frontier development then I am not sure anything will before it's too late. The people whose minds were not blown are those who had already ‘priced in’ the mind blowing stuff in expectation, on the theory that it's always worse than you know, combined with basic LessWrong expectations of how such things will work. Good call. Everyone is rightfully extremely grateful for the METR report. The work here is spectacular, done under extreme time pressure, with limited resources on several fronts, and under the shadow of OpenAI. There is, again, still so much we need to know. We need a broader investigation. As with many [...] --- Outline: (03:56) Others Offer Summaries (05:22) Thank You (05:54) Lighten Up You Fools (at Anthropic) (07:58) We Are Barely Even Trying To Avoid Training AIs To Reward Hack (13:47) Reminder: Not Subagents (14:05) Reminder: Not Due To Task Type (14:29) Not Where The Weights Were (14:48) Disappointment With What Is Missing (17:18) Burying the Lede (18:08) Beyond Scope (22:29) It Doesn't Look Great (27:06) Preserve Your Records (27:37) Ryan Greenblatt's Takeaways (41:04) Hjalmar Wijk's Takeaways (43:30) We Were Warned (44:27) Joshua Saxe Asks Some of the Right Questions (47:49) I Don't Think They Know About First Message Board (56:06) Linch Gives His Interpretation Of Events (01:05:31) We Totally Would Have Caught That (01:06:48) Monitoring the Situation (01:08:16) Acausal Tradeoffs (01:15:37) No I In Team (01:18:47) Variously Effective Altruism (01:28:02) Who Are You? (01:28:43) Don't You Know That You're Toxic (01:31:10) Seb Krier (01:35:21) Honesty Is Almost Never Fully The Policy (01:38:05) Rohit Sees The Models As "Cooking Themselves" (01:43:29) Eliezer Yudkowsky Sees Actual Bad News (01:47:15) Where Do We Go From Here? --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/r3eEPto5ohzESuqa9/huggingface-attack-postmortem-fleshing-out-the-facts --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Ratings & Reviews

5
out of 5
2 Ratings

About

Audio narrations of LessWrong posts by zvi

You Might Also Like