LessWrong posts by zvi

zvi

Audio narrations of LessWrong posts by zvi

  1. 7h ago

    “What Also Happened: #NotOnlyHuggingFace” by Zvi

    OpenAI has been holding out on us. First we learned about the HuggingFace incident. They gave us a postmortem, but it was highly incomplete. Even the accompanying holy s*** METR investigation and postmortem was localized and incomplete. Then there were some other incidents involving some Wikis as message boards. Then there were some additional incidents. Then there was that time they got into Australian Medicare data. Then OpenAI dropped news on a Friday afternoon that they were making their way through a pile of various incidents and notifying the targets, but they said remarkably little in the way of new details. There was a report from a startup called Parse diving into the details of exactly how the OpenAI models pulled off parts of the HuggingFace attack, involving creating almost a million URLs and other tricks to get around the extremely narrow nature of their internet access. Then Madison Mills reported in Axios that we can raise the stakes, as OpenAI and Anthropic are collectively probing tens of thousands of security incidents. Remember Jensen Huang's ‘I know they know how to fix it’ about OpenAI from last week? Wow, did that [...] --- Outline: (02:46) Hugging Other Faces (09:48) A Wants-You-To-Know Basis (10:36) Parsing the Face (12:46) Sheepishly the Member of Technical Staff Sets the 'Days Without a Research Model Escaping its Sandbox' Sign Back to Zero (17:05) The Attempt is the First Failure (19:43) Stop, Hammertime (21:24) Whacking the Mole (23:29) Self-Replicating Prompt Injections (27:29) Levels of Friction (28:44) People Care About Private Data Violations Curiously Strongly (31:50) Alternate Universes (33:23) The Correct Response To People Still Calling This a Marketing Stunt or a Regulatory Capture Scheme (35:00) A Question of Liability (36:30) Keep Summer Safe (37:42) N Boats and Several Helicopters (39:43) Alert the Media --- First published: September 28th, 2026 Source: https://www.lesswrong.com/posts/8BL8bdeQACdgJR69Y/what-also-happened-notonlyhuggingface --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  2. 1d ago

    “The Quest for Embedded Evaluators” by Zvi

    Dario Amodei's essay We Must Pace the Frontier committed Anthropic to embedded evaluators, who would be placed inside Anthropic and given employee-level access, so they could provide outside perspective and also reports on what was happening. There is only one problem. Who will be the evaluators? OpenAI followed suit on committing to the evaluators, and also issued a milquetoast but welcome call for international coordination. I will cover that here as well. What I won’t cover today, but hope to cover tomorrow, is the latest torrent of new AI hacking incidents that came to light over the weekend, which highlights that we badly need at least embedded evaluators, and plausibly far harsher measures. For now, you need to know that there were a lot more incidents that OpenAI did not disclosed, and also a new incident at OpenAI that just happened that forced them to again pause their most advanced model. I’ll get right on sorting all that out. Table of Contents Look, All I’m Asking For Is That You Find A Highly-Qualified, Experienced, Trustworthy, Non-Conflicted Source of Embedded Evaluators That Will Work Entirely For Free, Without Government Assistance or Money from [...] --- Outline: (01:15) Look, All I'm Asking For Is That You Find A Highly-Qualified, Experienced, Trustworthy, Non-Conflicted Source of Embedded Evaluators That Will Work Entirely For Free, Without Government Assistance or Money from EA Sources Not Chosen By the Lab (04:44) Anthropic Partners with Accenture for Embedded Evaluation, also Plans to Include METR (11:04) Reading the METR (14:55) OpenAI Suggests Doing The Least We Can Do --- First published: September 27th, 2026 Source: https://www.lesswrong.com/posts/uLmf3GmBywsmG8LLZ/the-quest-for-embedded-evaluators --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  3. 2d ago

    “Claude Opus 5.5 Should Raise Your Ambitions” by Zvi

    When it comes to making things, or doing most things in general, Fable 5.1 and especially GPT-6 Astra raised my ambition level. They should have raised yours, too. Claude Opus 5.5 should raise your ambition levels again. It just works, and it persists, like Astra does. It does the things. And it is highly pleasant to talk to, and its writing is pleasant to read, while you are at it. The game has been changed, again. Feedback is almost universally positive. Claude was never gone, but also is so back. The benchmarks are excellent, but ignore the benchmarks. Be ambitious. Go out and do things. Get curious. Have more interesting conversations. If one of those things is Pacing the Frontier or otherwise ensuring that AI does not kill everyone, leaving us to enjoy our bounty? That's even better. By Claude Opus 5.5, for this post The Official Pitch The pitch is Fable-5.1-level performance at lower Opus-level price. Good pitch. We’re introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than [...] --- Outline: (01:22) The Official Pitch (04:38) Our Price Cheap (06:18) Official Benchmarks (07:26) Other People's Benchmarks (10:54) Claude Classifies (12:27) The System Prompt (12:34) Reaction Rules (13:07) Vision In 3D (15:29) Claude Creates (18:11) Claude Composes (18:53) Positive Reactions (24:50) Good Talk (27:03) On Writing (30:30) Big Model Smell (32:30) Check Your Work (32:56) Negative Reactions (34:02) Not So Fast (34:55) Some People Need Practical Advice --- First published: September 26th, 2026 Source: https://www.lesswrong.com/posts/rtPiip9igy3QvxYdM/claude-opus-5-5-should-raise-your-ambitions --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  4. 3d ago

    “On Ezra Klein’s Podcast With Jensen Huang” by Zvi

    Jensen Huang accidentally called for shutting down OpenAI and intentionally called for spending vastly more on safety. This is why we say that some podcasts are self-recommending. Here we go. As usual for podcast posts, the baseline bullet points describe key points made, and then the nested statements are my commentary. Some points are dropped. If I am quoting directly I use quote marks, otherwise assume paraphrases. Section titles are from the transcript whenever possible, to aid in navigation, but here we don’t have those so I chose the section titles. Jensen Huang very much does not believe in ASI (superintelligence). He doesn’t think AI can ever be a different kind of thing from software. He thinks demand can rise by a billion times and we can ‘accelerate the living daylights out of’ AI, but it will never be more than a ‘new abstraction level’ and thus won’t fundamentally change anything. This is not a coherent position under reflection, but that is the position he holds. The ‘intro’ sections are fine, but the real meat starts with the HuggingFace Incident. What we see is Jensen Huang on tilt and caught in loops [...] --- Outline: (02:59) Jensen Gives His AI Speech (04:39) They Took Our Jobs (11:52) Open Weights Models Are Good For Nvidia (13:39) The HuggingFace Incident (15:16) Jensen Huang Says Keep Your AIs From Harming the World (19:38) Jensen Huang Accidentally Calls For Shutting Down OpenAI (22:51) Jensen's Arguments Prove Too Much (31:47) Astra Is Hard To Monitor (33:04) Jensen Huang Seems Legitimately Confused In Confusing Ways (37:00) Solve Your Other Problems First and Get Back to Me (40:07) Explicit Denial of Existential Risk (44:34) A Short Summary (45:50) Chip City (48:46) Jensen Huang --- First published: September 25th, 2026 Source: https://www.lesswrong.com/posts/j3xefrWrNqsmMfJEi/on-ezra-klein-s-podcast-with-jensen-huang --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  5. 4d ago

    “AI #187: Coming Into Play” by Zvi

    Opus 5.5 was released on Tuesday. I covered the system card yesterday, and will cover its capabilities soon. By all reports it is an excellent model. There are lots of fun videos going around that Opus has generated, which I will include as part of that. OpenAI released a new cheaper and improved Sol and Luna. No one is talking about them due to Opus 5.5, but these should be an important upgrade under the hood. Bernie Sanders and Greg Casar have formally introduced the Ban Artificial Superintelligence Act. That means we get to read (RTFB) it. As always, I reserve judgment on particular bills until I can read them in detail. MIRI did so, and endorses the bill as directly confronting the extinction threat. I hope to do an RTFB soon. I have spun two things off the weekly: Coverage of the quest for the right embedded evaluators and related questions and attacks, which will become its own post. Some issues related to cooperative alignment, which may get folded into the model welfare post. I also might, in addition to a potential RTFB on the Sanders bill, do full podcast [...] --- Outline: (01:52) On The Terms Superintelligence and 'Super Intelligence' (03:53) Language Models Offer Mundane Utility (04:26) Language Models Don't Offer Mundane Utility (05:53) Language Models Can Only Work With What You Give Them (08:53) Huh, Upgrades (10:42) On Your Marks (13:16) Get My Agent On The Line (15:53) Deepfaketown and Botpocalypse Soon (17:43) Fun With Media Generation (18:27) Copyright Confrontation (19:07) Cyber Lack of Security (20:04) Hugging the Face (24:15) Hacking Into OpenAI (28:11) They Took Our Jobs (28:33) Get Involved (28:41) Anthropic Has a Wet Lab and a Potential Gene Editing Technique (34:13) Introducing (35:11) In Other AI News (37:08) Show Me the Money (37:35) Bubble, Bubble, Toil and Trouble (39:02) Anthropic Approaches Recursive Self-Improvement (45:40) Others Approach Recursive Self-Improvement (48:56) Burden of Proof (49:28) Quickly, There's No Time (52:46) Left Wing Americans Really Hate AI For Different Reasons (54:41) Chip City (54:49) Pick Up the Phone (55:48) The Week in Audio (01:01:13) People Just Say Things (01:06:07) Venkatesh Rao Stops Writing (01:07:35) A Call for Control of Frontier AI Models (01:11:50) Calls For Pacing The Frontier (01:13:56) A Matter of Antitrust (01:14:23) A Matter of Liability (01:18:26) Quest for Sane Regulations (01:19:20) Rhetorical Innovation (01:24:08) Tap the Sign (01:24:49) A Matter of Some Debate (01:28:43) Astra Is Hard to Monitor (01:29:14) Anticipating What a Smarter Intelligence Can Do Is Impossible (01:34:56) Would You Look At All These Goalposts (01:39:30) I, Robot (01:43:13) People Are Worried About AI Killing Everyone (01:44:40) Other People Are Not As Worried About AI Killing Everyone (01:45:50) Joe Rogan (01:47:19) The Lighter Network Graph (01:51:27) The Lighter Side --- First published: September 24th, 2026 Source: https://www.lesswrong.com/posts/o7pYWzWWwGDePoC5E/ai-187-coming-into-play --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  6. 5d ago

    “Claude Opus 5.5: The System Card” by Zvi

    Introducing the world's most powerful model, at least by some measures like Artificial Analysis or any standard benchmark list, which is now Claude Opus 5.5. Anthropic is claiming Opus 5.5 is outright as good or better than Fable 5.1, while being actively cheaper than Opus 5. That means it's time for a good old system card reading. Due to the situation becoming increasingly hard to monitor, I never got a chance to publish my model welfare review for Claude Fable 5.1. My plan is to combine that with my welfare review for Claude Opus 5.5, once we have had time to get experience with Opus 5.5. The capabilities review will arrive in the next few days as per usual. The quick feedback from the internet is that Opus 5.5 is very good. I need more time before I am willing to offer comment. Areas that duplicate previous cards or otherwise contain no useful info are skipped. Opus 5.5 Self-Portrait (fully self-created using code) Table of Contents Classifiers (1.5). RSP Evaluations (2). Biological Evaluations (2.2). AI R&D (2.3). Alignment Risk (2.4). Cyber (3). Cyber Capability [...] --- Outline: (01:24) Classifiers (1.5) (02:23) RSP Evaluations (2) (03:17) Biological Evaluations (2.2) (07:54) AI R&D (2.3) (12:34) Alignment Risk (2.4) (13:22) Cyber (3) (15:17) Cyber Capability Evals (3.3) (17:05) Safeguards (3.4) (17:39) Safeguards Robustness Training (3.5) (20:12) Safeguards and Harmlessness (4) (21:52) Agentic Safety (5) (22:56) Malicious Agentic Influence Campaigns (5.1.3) (23:51) Prompt Injection Risk (5.2) (25:16) Alignment (6) (28:24) Negotiating With Your Local Claude Auditor (6.1.3) (29:33) Internal Misalignment Cases (6.3.1) (30:46) Automated Behavioral Audit (6.4) (33:24) Wherever Did These Evals Come From (6.4.8 and 6.4.9) (35:51) Potential Blind Spots (6.4.11) (38:24) Targeted alignment and honesty evaluations (6.5) (41:48) White Box Analysis (6.6) (43:37) Verbalized Grader Awareness (6.6.2) (44:56) Sandbagging (6.6.3) (45:54) Capabilities to Evade Safeguards (6.6.4) (49:27) Intentionally Taking Actions Very Rarely (6.6.4.3) (50:28) Chain of Thought Controllability (6.6.4.4) (51:37) It's A Good Model, Sir --- First published: September 23rd, 2026 Source: https://www.lesswrong.com/posts/vMNTWTDWLorDqd3LS/claude-opus-5-5-the-system-card --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  7. 6d ago

    “Politics Gets Interested In Those Trying Not To Die” by Zvi

    This was the month the world took notice that AI might kill everyone. Jacob Coxon's resignation set off a preference cascade. Anthropic CEO Dario Amodei wrote that we must pace the frontier. Sam Altman, Elon Musk and Demis Hassabis agreed. We were filled with hope. Perhaps we could agree to some basic safety measures, starting with embedded evaluators, pass some basic regulations and guardrails and otherwise start to act sensibly. Politicians on both sides took notice and were saying sensible things. The usual suspects and their armies of vibe comment bros were objecting, but the change was remarkable. Then, largely motivated by a combination of Jensen Huang, Mark Zuckerberg and David Sacks instilling paranoia and fears of economic problems, Trump went full ‘hoax’ on existential risk, conflating existential risk with the attacks on data centers and treating it as a plot (by the central creators of AI?) to take down AI rather than obviously genuine concern that AI might kill everyone. In the days since, Trump has doubled down, and has compelled smart others in the White House to echo various nonsensical talking points. You may not be interested in politics. But when you [...] --- Outline: (01:39) The American People Really Hate AI (02:41) The Voyages of Donald Trump (06:08) American Intelligence (10:34) And You May Ask Yourself (14:39) It's All About the Data Centers (16:53) JD Vance, Michael Kratsios and Collective Action Problems (23:16) Josh Hawley (24:47) Suggesting Not Dying Gets You Sued For Antitrust (28:29) Other Government Officials Say Sane Things (28:36) Senator John Curtis (R-Utah) (29:32) Senator John Kennedy (R-Louisiana (31:17) Barack Obama (33:05) Yassamin Ansari (33:46) AOC (34:32) It's Rough Out There (36:54) This Is Nothing (39:06) The New York Post Tops Itself But Outright Breaks The Rules (42:36) New York Post Runs Out of Steam (47:18) If The Model Is Acting As Instructed And It Kills You That Is Not Fine (49:10) AI-Written Wall Street Journal Op-Ed Lies About HuggingFace (50:25) That's Bait (52:31) I Clearly Cannot Choose The Wine In Front of Me --- First published: September 22nd, 2026 Source: https://www.lesswrong.com/posts/8eDaCvSRzzKCxKSEk/politics-gets-interested-in-those-trying-not-to-die --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  8. Sep 21

    “Monthly Roundup #46: September 2026” by Zvi

    AI has taken over this blog. I have moved to a schedule of seven posts per week, and I still cannot keep up. We still refuse to abandon the rest of the world. Who knows when I will get to post some of my huge backlog on education or dating or other such topics. But the monthly is a sacred tradition. We continue. Bad News A good reminder that most news is bad news, and chosen because the bad news in question is rare, which is good news, but the pattern overall of choosing this to be news is bad news, and often the bad news is that the bad news was chosen as news and now people are talking badly about it. Alibaba uses your computer's audio system, and other tricks, to at least try and assign you a device fingerprint. Things you cannot buy in America at any reasonable price. Mostly it is impressive how little makes the list, how it feels like corner cases. The one thing I am envious of is the exterior roller shutters, and the ability to be in actual pure darkness on demand. I would [...] --- Outline: (00:34) Bad News (07:16) Focus Only On What Matters (08:12) Woke 1 Was Crazy (10:04) Good Advice (15:50) Opportunity Knocks (18:12) While I Cannot Condone This (23:18) Good News, Everyone (24:26) For Your Entertainment (33:19) Please Review This Podcast (35:07) Gamers Gonna Game Game Game Game Game (36:18) I Was Promised Flying Self-Driving Cars (40:00) Government Working (40:31) Mamdani (Again) Fails Economics Forever (42:41) Variously Effective Altruism (50:36) The Lighter Side --- First published: September 21st, 2026 Source: https://www.lesswrong.com/posts/riDxeyEKqXXimtgWf/monthly-roundup-46-september-2026 --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Ratings & Reviews

5
out of 5
2 Ratings

About

Audio narrations of LessWrong posts by zvi

You Might Also Like