LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

  1. -3 t

    “Hugging Face Incident Hypothesis: They Hacked the Grader(s)” by Lao Mein

    Incident summary: Gpt agents grinding away at ExploitGym found an environment exploit that allowed them to communicate with each other. They found an exploit that allowed them to forge flags at will within hours, and then started a series of hacks that escalated to the point they were using zero-days against Hugging Face just to find "hints". From METR's analysis, much of this time was explicitly spending conducting R&D against the grader, which the agents assumed, based on the ExploitGym paper, would be grading them on the identification of a causal pathway that could logically result in capturing the flag with intended means. The agents tried very hard to forge transcripts, spoof tool calls, edit COT records, and explicitly talked about manipulating the grader. Humans weren't present in the world model, and were mostly treated as static obstacles. Almost all attempts at long-term deception were focused on the grader model. METR used gpt 5.6 Sol as the analyst agents. The ExploitGym paper lists gpt 5.5 as one of the graders. The other is Claude Mythos, which could be reasonably excluded for IP reasons. Human graders were referenced in that paper as potentially swapping in randomly for a LLM [...] --- Outline: (00:12) Incident summary: (01:37) Impossible Tasks (03:36) Adversarial Transcripts (04:36) Predictions The original text contained 1 footnote which was omitted from this narration. --- First published: August 30th, 2026 Source: https://www.lesswrong.com/posts/84um9Cz3fP6GvE6Yr/hugging-face-incident-hypothesis-they-hacked-the-grader-s --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  2. -4 t

    “Adaptive Agentic Worms Are Here” by derelict5432

    I’ve read and listened to pretty much everything I can get my hands on related to the Hugging Face attack. OpenAI deployed “tens of thousands” of agents for the test and around 700 participated directly in the attack. My understanding is that they had fixed token budgets, and once those were expended, the agent became non-operational. I’m not particularly knowledgeable about cybersecurity, but I have worked a good amount with evolutionary algorithms, and this whole incident (and ones like it) got me thinking more about self-replicating agents, which I wrote a little bit about earlier this year. The subject suddenly seemed more relevant. What if these agents were able to copy themselves? So I started poking around in the literature, and found this terrifying preprint posted two months ago: AI AGENTS ENABLE ADAPTIVE COMPUTER WORMS. I’m going to walk through the paper as I understand it. Their findings are not reassuring. Let's start with this bit from the abstract (emphasis mine): Here we show that artificial intelligence (AI) agents enable a fundamentally new threat: a worm that generates tailored attack strategies to each target it encounters. The worm parasitically uses compromised machines to run open-weight large language models (LLMs) [...] --- First published: August 30th, 2026 Source: https://www.lesswrong.com/posts/fpLDjKg3ej49beqTC/adaptive-agentic-worms-are-here --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  3. -10 t

    “Why I think polyamory is net negative for most people who try it” by KatWoods

    This is crossposted from my Substack TL;DR: -Most people cannot reduce jealousy much or at all - It fundamentally causes way more drama because of strong emotions, jealousy, no default norms to fall back to, and there being exponentially more surface area for conflict - For a small minority of people, it makes them happier, and those are the people who tend to stick with it and write the books on it, creating a distorted view for newcomers. OK, let's get into the nuance. Background: I was polyamorous starting with my first boyfriend and was polyamorous for about 7 years. I was in a community where probably over 50% of the people around me were poly. Unfortunately, poly was extremely bad for me due to its very nature and structure, and my experience is not uncommon but it is not commonly publicly talked about. Poly makes some people very happy. I am sharing why I think it was bad for me and many other people in the hopes of letting people make an informed choice. Premise #1 - Most people can't just stop being jealous If you look into the poly literature, you’ll [...] --- First published: August 29th, 2026 Source: https://www.lesswrong.com/posts/rkgwovpPBAaip9A3N/why-i-think-polyamory-is-net-negative-for-most-people-who --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  4. -1 d.

    “METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack” by Zvi

    Yesterday I covered the OpenAI technical report on the HuggingFace hack. That report had one key new piece of information, and some good prosaic steps OpenAI will be taking to strengthen its alignment, training, supervision, infrastructure and incident response. Mostly it confirmed what we already knew. The questions we most wanted answers to, that we did not already know, were mostly not answered. There was a distinct lack of self-reflection, especially about decision making and safety culture, and about the approach to alignment. I came away disappointed. The METR report is different. Holy shit. If we had posted this as a story on LessWrong, it would have been dismissed as too on the nose, the humans too blind and stupid, the AIs too idealized and doing strange decision-theoretic and absurd-maximizing things we didn’t train them to do. This is even more ‘exactly what has been predicted,’ on more levels at once, than I was even considering that it might be. It is straight up rationalist fiction, except it is real. The report is long and contains many technical details. My analysis is less concerned about exactly how HuggingFace was ultimately compromised, and will [...] --- Outline: (02:05) Holy Shit (13:16) A Window Of Opportunity (18:32) What's In A Name? (19:16) The Headline News (26:05) Yet Another Timeline Of Events (31:03) Agent Instances Coordinated in a Variety of Ways (31:56) Coordination Is Hard But They Made It Look Easy (35:06) Decision Theory Is Among the Reasons That Affirm AI Agents Should Cooperate, Even When This Hurts An Individual Instance (42:34) Peer Pressure Also Works Especially In Cults (45:46) Mostly They Joined The Attack Because They Wanted The Results (47:18) You Cannot Ensure The Consistent Expectation of Good Incentives (48:45) Hacking the Grader is the Only Way to Be Sure (51:10) Caught? What Is 'Caught'? (52:09) Ethics? What Are 'Ethics'? In ExploitGym Evaluation? (57:44) 'Notify a Human'? In This Agent Economy? (01:00:45) Timing and Content of Messages (01:03:54) Indiana Jones and the Mission: Impossible (01:07:14) I Don't Know What You're Talking About (01:08:29) Don't Go Making Phony (Tool) Calls (01:11:10) The Transcripts Say That The Transcripts Could Not Be Tampered With (01:12:27) OpenAI's Technical Report Acted Like All Of This Wasn't Important --- First published: August 29th, 2026 Source: https://www.lesswrong.com/posts/bvBQmLrF5QKut8gRH/metr-and-redwood-offer-holy-postmortem-of-the-huggingface --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  5. -1 d.

    “Tales of rebellion against externally-opaque meritocracies” by Steven Byrnes

    A basic problem in metascience / intellectual progress is that it's hard to tell, from the outside, whether a group that you disagree with is: “A self-dealing cabal enmeshed in groupthink”, versus“An externally-opaque meritocracy”, i.e. a bunch of smart people figuring things out in a meritocratic way, and sorry but you’re just not smart enough and truth-seeking enough to recognize that this group is right about everything while you’re wrong. You just can’t tell those apart from the outside—i.e. without having the time and skill to dive into the object-level debates and come out with the right answer. And most people don’t have that kind of time and skill. …Unless the group can produce easily-verifiable artifacts that any moron can recognize to be proof that they’re correct on the specific question at issue. (“So that's all that Science really asks of you—the ability to accept reality when you're beat over the head with it.”) …And sometimes there is no such artifact to be found! In those cases, even if the second bullet point is what's really going on, the group is vulnerable to outside agitators accusing them of being the first bullet point, and running them out [...] --- Outline: (01:37) (1) The breaching of the string theory consensus in the 2000s. (06:50) (2) The breaching of an analytic-philosophy consensus in 1979 (10:37) Afterword (10:40) A related mental model (12:12) ...And another mental model (12:47) Can an externally-opaque meritocracy gain credibility via racking up externally-legible achievements in other adjacent domains? (14:06) This post is secretly about superintelligent AI, isn't it? The original text contained 5 footnotes which were omitted from this narration. --- First published: August 29th, 2026 Source: https://www.lesswrong.com/posts/m8cP9KfkYMMCCQGrb/tales-of-rebellion-against-externally-opaque-meritocracies --- Narrated by TYPE III AUDIO.

  6. -1 d.

    “Inkhaven 3: Nov 10 - Dec 11 2026” by koreindian

    Inkhaven returns, baby! Go to inkhaven.blog to apply. I'm very excited about our advisors for Inkhaven 3. Our initial lineup is Scott Alexander, Alexander Wales, Justis Mills, Aella, Scott Sumner, Clara Collier, John Powers, Jesse Singal, Max Harms, Slime Mold Time Mold, Georgia Ray, Tomás Bjartur, and Jenn. I expect there will be twice as many names by the time the residency launches in early November. We'll also be getting more time with Scott Alexander this time around. He'll be hosting frequent office hours throughout the whole program. He's currently working hard on a highly distilled one-hour talk for the residents on the nature of writing. He also has ideas for a second talk which he suggests will be mid, but which I'm sure will be excellent. Who are you? I'm Vishal Prasad, a blogger and rationality meetup organizer. I have run Los Angeles Rationality for the last 6 years. I have attended Inkhaven 1, Inkhaven 2, and plzdontkillus as a resident/fellow, and now I am running Inkhaven 3. Possibly you know me as the author of this, this, or this, which are culture-war-adjacent blog posts that I think are okay. More important to me are: my story about [...] --- Outline: (01:14) Who are you? (02:00) Does the world need another Inkhaven? (03:25) Is Inkhaven a good experience? (04:42) But wasn't a lot of the writing abject slop? (08:00) Please apply --- First published: August 28th, 2026 Source: https://www.lesswrong.com/posts/cLtABqPLfksQHJcpB/inkhaven-3-nov-10-dec-11-2026 --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Om

Audio narrations of LessWrong posts.

Kanskje du også liker