LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

  1. 35 min fa

    “Stop doing decision theory without metaphysics” by Elias Schmied

    [Epistemic status: rant] There's something that annoys me about the reoccurring debates on decision theory in this corner of the internet. Take a simple blackmail scenario: Omega, a near-perfect predictor, knows a piece of embarrassing information about you. He threatens you that he will release it to the public if you don’t pay him 100$. However, making the threat is slightly costly to him, and he wouldn’t have done it if he hadn’t predicted that you would pay. Do you pay? Let's say we want to argue for the Functional Decision Theory (FDT) answer that you shouldn’t pay. It was originally motivated by the observation that such agents seem to achieve higher utility (“rationality is about winning”), since they don’t get blackmailed in the first place. But that only leads to making yourself into such an agent in advance (which everybody generally agrees you should do[1]) - it doesn’t clearly apply when you are already being blackmailed and have never thought about the question before, or if you are an AI who was just created and is instantly blackmailed before being able to self-modify or make precommitments. I see three broad ways to make FDT's recommendation make sense in [...] The original text contained 13 footnotes which were omitted from this narration. --- First published: July 20th, 2026 Source: https://www.lesswrong.com/posts/oZzRHiSZPcjrWHeoE/stop-doing-decision-theory-without-metaphysics --- Narrated by TYPE III AUDIO.

  2. 49 min fa

    “War – What is it Good For?” by kqr

    Spoiler: positive expected utility. That's what it's good for. In The War Trap[1], Bruce Bueno de Mesquita lays out a theory of war that beats all other theories. This is a great example of a model that is parsimonious, wrong, and useful. For every assumption Bueno de Mesquita makes, the reader goes, “What? That's not how things work.” Yet when we measure the things Bueno de Mesquita asks us to measure, and apply the equations he lays out, the result ends up closely matching reality. What's better, his equations arrive at the same results that previous ad hoc theories did – and also explains observations those theories could not account for. The work Bueno de Mesquita did on The War Trap is creative, intelligent, and inspiring. We don’t know what a potential war will be like Here's the scene for the book: there's an active international dispute, and one country has issued a threat of solving it with violence. We are faced with the same questions that troubled Tolstoy when he wrote War and Peace: Will the threat escalate into a full-on war?If so, which countries will join the war?How bloody will the war be? Tolstoy's [...] --- Outline: (00:57) We don't know what a potential war will be like (02:05) We measure military alliances and fuel usage (03:52) Then we compute expected utilities (06:20) Predictions from expected utility calculus (08:25) The surprising effects of good relations (10:08) Predicting the severity of wars (12:31) It's impossible until the right person tries The original text contained 9 footnotes which were omitted from this narration. --- First published: July 20th, 2026 Source: https://www.lesswrong.com/posts/qnDvrbNz5eteax3cN/war-what-is-it-good-for --- Narrated by TYPE III AUDIO.

  3. 3 h fa

    “Against the AI framing multiverse: Introducing AI StopWatch” by tanagrabeast

    In my long years as a classroom teacher, it was my experience that the kid most likely to speak up during discussion was the one who did the reading. I think it's true for adults, too. I know that's not exactly revelatory, but it's one of the guiding principles behind AI StopWatch, the experimental newsroom we (parts of the MIRI comms team) launched in May, after a month of closed beta testing. StopWatch's other guiding principle is that the reading needs to make sense. Imagine if every page that kid read came from a different book by the same name, and if most of these different books were just reflections of what people who didn’t read it imagined it would be. If the book in question were To Kill a Mockingbird, then (speaking from experience) this would mean that on one page, the Finch family is Black and oppressed. The next page is a hunting manual. Flip the page again, and Scout is a boy. Flip to a page near the end, and Atticus might win the case. This is how the media landscape around AI looks to me. Even when the facts agree, the frames are so varied [...] --- First published: July 20th, 2026 Source: https://www.lesswrong.com/posts/Rk57ePmRsw4C5PLm2/against-the-ai-framing-multiverse-introducing-ai-stopwatch --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  4. 3 h fa

    “We’re talking past our models; or, How a model defined its “evil” vector as dread” by jcksanderson

    Summary We train a new token—a neologism (Hewitt et al.)—for a model, but unlike Hewitt et al., we train it on data the model generated while steered with a persona vector.To learn how the model interprets this steering vector, we then ask the model to a) respond in the style of this neologism, and b) explain it.Responses generated with the neologism are substantially more similar to the steering vector (larger projection values) than responses generated with the steering vector itself, while being more coherent and trait-expressive (per an LLM judge).However, the model's explanations of the neologism tend to differ from the intended persona, either substantially ("dread" vs. the intended "evil") or subtly ("warmth" vs. "sycophancy").Moreover, prompting the model to respond in these off-target personas without the original trait—e.g. "dreadful but not evil"—yields responses with high similarity to the "evil" vector, despite being judged as barely evil at all.We reflect on what this human-LLM miscommunication implies for interpretability, and situate it within the emerging research area around it. Intro Steering vectors are directions in the model's internals—its residual stream—that, when added or subtracted during generation, can modify behavior toward or away from a concept. A large body [...] --- Outline: (00:13) Summary (01:29) Intro (02:09) Generating the steering vectors (03:13) Do the steering vectors work? (05:27) Neologisms (10:15) Q&A (11:31) Misgeneralization (14:35) What makes neologisms so effective? (16:08) The Whole Point is Miscommunication (18:09) So what should we do? The original text contained 10 footnotes which were omitted from this narration. --- First published: July 20th, 2026 Source: https://www.lesswrong.com/posts/ktCYxLgdtFR2fDw7J/we-re-talking-past-our-models-or-how-a-model-defined-its --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  5. 21 h fa

    “Many alignment techniques work by training one model and deploying another” by cloud

    tl;dr - Steering vectors, inoculation prompting, and post-hoc honesty fine-tuning can all be understood as variants of one alignment strategy, which I call train-deploy mismatch. Each trains the model in one configuration and deploys it in another. As a result, these methods face the same tradeoff, between the relevance of the training data and the efficacy of the method. Note: Others have had similar ideas and shaped my thinking here including Sam Marks, Ariana Azarbal, Victor Gillioz, Alex Turner, Jacob Goldman-Wetzler, Jake Mendel, Daniel Tan, and Fabien Roger. Thanks to Monte MacDiarmid, Nat McAleese, Shawn Hu, and Jake Ward for input on an earlier draft. Background AI alignment is hard largely because we don't know how to specify what we want. Instead, we train models on proxies for what we want: labels and reward functions defined on data distributions chosen such that we hope the model will perform as desired when deployed into the world. This approach has worked well so far, but given increasing model capabilities, it may stop working— models may misgeneralize their training to catastrophically bad behavior in deployment. A pressing open problem is to figure out how to get models to generalize the properties that [...] --- Outline: (00:55) Background (02:06) Train-deploy mismatch as a general alignment strategy (04:17) Fundamental tradeoffs (07:01) Sidebar: is everything train-deploy mismatch? (07:43) Open threads (11:11) Closing thoughts The original text contained 1 footnote which was omitted from this narration. --- First published: July 19th, 2026 Source: https://www.lesswrong.com/posts/syAbdNei8BWeP2RPo/many-alignment-techniques-work-by-training-one-model-and --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  6. 1 g fa

    “A Post-Mortem for My Goal Crystallisation Project” by atryt0ne, Jason R Brown

    I concluded my MARS 4.0 project titled 'Goal Crystallisation' with Anaïs Berkes and Lukas Gebhard under the mentorship of @Cameron Tice and @Jason Brown. We wanted to find out how important a threat scheming was. In particular, we wanted to find out whether a perfectly alignment faking agent could preserve its hidden agenda as it receives constant positive feedback for complying with requests. How does a hidden goal evolve over time as an alignment faking agent is subjected to SFT and RL training? Unfortunately, we were unable to reach a positive result and had to shut the project down. This post discusses our failure points and where we could have possibly found an earlier signal to stop. In the first section, I talk about our plan for this project.In the second section, I break down the key issues that led to our failure.In the last section, I give some advice for people interested in doing a project on this topic in the future. The Plan Figure 2 from the original alignment faking paper In Figure 2 of the original alignment faking paper, we see that the RL training increased the rate at which the model answers harmful queries in [...] --- Outline: (01:11) The Plan (03:37) Why We Failed (04:35) Issue 1 (06:14) Issue 2 (11:08) Issue 3 (12:57) The Final Nail in the Coffin (13:45) So You Want To Study Goal Crystallisation? (17:00) Conclusion --- First published: July 17th, 2026 Source: https://www.lesswrong.com/posts/BGh5qYo5XTLFXqjDo/a-post-mortem-for-my-goal-crystallisation-project --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  7. 1 g fa

    “Nuances in the Workings of the Eye and Retina” by Hieronym, Julian Bradshaw

    Image edited from source at Britannica The eye (any eye) is a miracle of evolution, which has filled its structure at every scale with an apparent intentionality that is the hallmark of complicated systems under intensive selective pressure. It is very much possible to look at almost every design feature and state a compelling reason why it's there, which makes it a rich place for the engineering-minded to find fun little design features. This post is meant to discuss topics that are less frequently discussed elsewhere, and not often discussed in introductory courses, rather than anything that can be quickly gleaned off of common summaries. Another way to describe it is that I intend to dive a bit into the kind of fridge-logic questions you might come up with randomly after opening the evolutionary eye-fridge, particularly the ones that turn out to have interesting answers. As such, this post is going to be a bit of a grab bag of topics, rather than a focused point-by-point breakdown of the eye. I’m still going to group topics broadly into general categories, just for organizational reasons. One thing I will not go into is visual processing in the retina or visual [...] --- Outline: (01:35) Part I: Optics and the Lens (01:40) Biology Basics (Brief) (03:04) Chromatic Aberration (refractive lens drawbacks) (04:52) Why Everything Looks Weird Underwater (but not for fish) (07:37) Crystallins: Building Large Functional Structures Biologically (spoiler: it's frequently dead cells) (10:55) Failure Modes (practical knowledge that may apply to you!) (11:20) Myopia and Hyperopia (12:04) Presbyopia (12:38) Cataracts (13:26) Glaucoma (15:07) Macular Degeneration (17:31) Diabetes-Related Retinopathy (17:50) Retinal Detachment (18:50) Part II: Photoreception (18:55) Basics II: Rods, Cones, and Color as a Matter of Information (22:49) Detecting Photons with Cells and Chemicals (quantum mechanics!) (27:52) The Macula and Fovea (efficiency compromises) (29:25) That Thing About the Retina Being Backwards (also cephalopods) (32:14) Tapetum Lucidum (glowy cat eyes) (32:58) Infrared and UV vision (visible light range isn't that arbitrary, also snakes) (35:57) Saccades (more efficiency tricks) (38:01) Teaser - Visual Processing in the Retina (38:43) Summary The original text contained 38 footnotes which were omitted from this narration. --- First published: July 18th, 2026 Source: https://www.lesswrong.com/posts/zvpocS7CswgT8bw7E/nuances-in-the-workings-of-the-eye-and-retina --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Descrizione

Audio narrations of LessWrong posts.

Potrebbero piacerti anche…