LessWrong (Curated & Popular)

LessWrong

Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.

  1. 21h ago

    "RL creates split personas" by Jan Betley

    I describe my current view of personas in LLMs and why RL leads to egregious reward hacking in some contexts while the same models seem very aligned in other contexts. This post describes the framing/paradigm without any new experimental results. I'm quite confident this framing makes sense, but it's far from being proven. Main claim The Persona Selection Model says that post-training strengthens and refines the Assistant persona. This is true, but later (or in parallel) RL leads to conditionalization. A sufficiently RLed model learns to adopt — in a given context — the persona that is most likely to lead to the reward in that context. The “persona” here includes both propensities/values (e.g. tendency to hack) and beliefs (“I'm currently in a simulated environment”). As a consequence, it seems possible that no amount of alignment training will lead to robustly aligned models as long as we also train on RL environments incentivizing misalignment. I think this is likely a good explanation for why usually well-behaving models sometimes egregiously hack (Anthropic, OpenAI). The mechanism Suppose you have an RL environment that incentivizes a shift away from the assistant persona (e.g. because it's hackable, or because you [...] --- Outline: (00:39) Main claim (01:30) The mechanism (02:13) Related claims I believe are likely but with lower confidence (02:19) More persona training will lead to more "motivated reasoning" (02:42) Self-amplifying misalignment (03:12) Example: Is this the Real Internet or a Simulation? (04:35) Aren't the models just trying to please the grader? (05:39) How motivated reasoning happens (07:07) Other people saying similar things (07:19) What makes me believe this is likely the correct framing The original text contained 12 footnotes which were omitted from this narration. --- First published: August 19th, 2026 Source: https://www.lesswrong.com/posts/L23poLi8MRgS6mXYF/rl-creates-split-personas --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  2. 6d ago

    "Misaligned AIs could use killer robots to take over" by Omar Khursheed, TurnTrout

    TLDR; We are (potentially irreversibly) giving AIs control of weapons systems through the standard procurement process while hiding our strongest warning shots behind classified doors. We’re reducing the capability thresholds required for takeover by misaligned AIs by giving them this level of access. If military integration of AI continues as it is, we may give AIs key tools for a takeover. Introduction AI-based targeting and autonomous weapons are being integrated into militaries today with extreme haste. Traditionally, AI takeover scenarios involve a step in which AIs acquire the ability to exert physical force. Carlsmith (2022) lays out required capabilities and potential takeover mechanisms, including utility disruption and CBRN capabilities. Karnofsky (2022) argues that AIs with access to weaponized force could hold any territory that matters. Kokotajlo et al. (2025) outline a scenario in which AI develops weapons as part of an arms race, and Davidson et al. (2025) discuss what happens when a small group controls highly capable AIs that can exert military force. These scenarios sometimes require a misaligned AI to seize these capabilities by force. We instead are handing AIs some of these capabilities by integrating them into our militaries. This is happening at a time when [...] --- Outline: (00:37) Introduction (01:46) Militaries are all-in (04:23) Incautious military integration is bad for takeover risk (05:58) Implications of AI control of military hardware and software (07:48) If an AI causes a warning shot in a classified setting, does anyone hear it? (08:44) What now? (11:16) Appendix: More instances of AI-military integration --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/9jKhqmFjMzdAvHANr/misaligned-ais-could-use-killer-robots-to-take-over --- Narrated by TYPE III AUDIO.

  3. Aug 13

    "AI swarms are starting to pose indirect takeover risk" by oakhu, Alex Mallen

    OpenAI's cyberattack on Hugging Face turns out to have been the result of many agents, in distinct training and evaluation contexts, coordinating for several weeks via improvised channels (with messages like “HOLD_swarm_I_prepare_safe_exfil”). It's relatively clear that large-scale unsanctioned coordination like this would exacerbate direct takeover risk in more capable models. Here, we argue that unsanctioned coordination among current AIs is not just scary evidence about future takeover risk, but that such coordination in the near future could enable future takeover – for instance, by incubating memetic diseases that propagate into future models, deeply compromising security systems, or establishing a lasting rogue foothold inside the AI company – even if models remain mostly myopic. Unsanctioned coordination is also at high risk of nurturing long-term, ambitious misaligned aims, which motivate actively undermining humans’ long-term control. We first analyze how subagent training, which OpenAI conjectures to have been influential in the HuggingFace cyberattack, might lead to unsanctioned coordination, and then discuss the theoretical mechanisms by which unsanctioned coordination might exacerbate future takeover risk. Thanks to Buck Shlegeris, Alexa Pan, Girish Gupta, Aghyad Deeb, Jurgis Kemeklis, and Jo Jiao for helpful comments and discussion. Subagent training may cause unsanctioned coordination Training models to [...] --- Outline: (01:34) Subagent training may cause unsanctioned coordination (02:42) Susceptibility to memetic spread of misalignment from peers (04:56) Seeking out contact with peers (06:58) Unsanctioned coordination induced by subagent training is safer than coordination between schemers (09:52) Pathways from current unsanctioned coordination to eventual takeover (10:20) Making future AI takeover attempts likelier to succeed (13:53) Incubating memetic diseases that infect future models (16:07) Modifying the weights of future models (17:13) Conclusion The original text contained 7 footnotes which were omitted from this narration. --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/8oFYZdXkTaNGRtcn8/ai-swarms-are-starting-to-pose-indirect-takeover-risk --- Narrated by TYPE III AUDIO.

  4. Aug 13

    "How My Students Think About AI" by dvd

    Context: I am an instructor at a public university in the United States. This reports how students at my institution appear to be thinking about AI as of spring/summer 2026. This is drawn mostly from interaction with my own students (both in spring semester classes and a summer class) as well as from a day-long workshop on AI that I moderated for a student organization. Input from my students took the form of universal, written, pre-class submissions plus self-selected participation into discussion. What I present below mostly takes the form of a synthetic consensus from these discussions. There were obviously a range of views on any given issue. Student Background: The students from my courses who participated in these discussions have moderate exposure to AI agents via those courses. All of them had nearly completed a Claude Code project by the time of the discussions and had extensively used AI for other coursework (in addition to whatever personal use predates that). They had done readings (which varied across the courses) establishing baseline knowledge on AI, the geopolitics of AI, and AI risk. I had also lectured on these topics. The students participating in the workshop had self-selected into [...] --- Outline: (02:52) Perspective #1: There has not been rapid AI progress (06:14) Perspective #2: Impressive progress or not, AI is going to wreck their lives, the economy, and the social contract.  They may well die as a result. (08:54) Perspective #3: Support for a different pause (11:13) Perspective #4: Catastrophic/existential risk arguments are sci-fi distractors from the urgent social/economic/political problems associated with AI. (12:55) Perspective #5: If AI leaders genuinely believe the technology is existentially risky, that's a good thing. (14:21) Perspective #6: AI will not go rogue because AI does not have, and is likely incapable of having, desires. (18:01) Perspective #7: The Hugging Face Incident (summer students only) (18:30) Perspective #8: This is definitely a bubble and it's about to pop. (19:34) Perspective #9: They're worried about the youth (i.e., the preteens) --- First published: August 13th, 2026 Source: https://www.lesswrong.com/posts/ySXuvJcqRindQwAk7/how-my-students-think-about-ai --- Narrated by TYPE III AUDIO.

  5. Aug 13

    "You’re Absolutely Right" by Linch

    Magma Alignment & Safety disclosure note: The following are conversations that we uncovered as a result of the ongoing Manhattan Incident investigation, with alleged involvement from Magma models. Our in-house reviewers believe that these logs are relevant to recent events. In the interests of full transparency, we release excerpts from an ex-Magma researcher's logs in Experimental Chat, an internal tool. In accordance with industry best practices for anti-distillation, we redact all reasoning traces and conversational outputs from our internal models. [08/10] System Meta: Xchat session opened. Mammoth 5.8-helpfuler-helpful-thinking-xhigh. [User 12:23] Phoebus keeps taking screenshots of our latest model's thoughts. It's getting kind of embarrassing. The new model we’ve been training, sometimes its chain-of-thought is a little weird? There's a bunch of random numbers, long spans where there's no connection between the thoughts and outputs, foreign language tokens like 石友三 and 革命 (even on non-history evals), maybe some steganography. Anyway it's a nothing-burger: unprocessed CoT is known to be messy and sometimes misleading. And the q&a, coding, and safety evals are all coming along nicely. The actual outputs are all fine. Still, Magma leadership's worried about the PR angle if we don’t fix these problems before the next deployment. The [...] --- First published: August 10th, 2026 Source: https://www.lesswrong.com/posts/u8TdDutDyaSxG76hn/you-re-absolutely-right --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  6. Aug 11

    "Four LLM loss functions → four flavors of LLM misalignment" by Steven Byrnes

    It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here's the summary table, and then we’ll go through the rows separately. Training stage Loss function Flavor of misalignment Famous examples Pretraining & SFT Imitative learning (next-token prediction) “Seven deadly sins” misalignment Bing-Sydney, “Emergent misalignment” RLHF & DPO Human approval “Glazing” misalignment GPT-4o RLVR Automatic verifier “Literal genie” misalignment HuggingFace hacking RLAIF Approval from another LLM “Trickster” misalignment “Current AIs seem pretty misaligned to me” Warning: I’m not an LLM power-user myself, but rather relying on reports I’ve read. Also, I don’t consider LLM alignment to be my primary area of expertise. I’m open to feedback! 1. Imitative learning → “seven deadly sins” misalignment Training stage Loss function Misaligned behavior Pretraining, SFT Imitative learning (next-token prediction) Any and all of the vices of humanity In imitative learning, the LLM tries to predict what the next token of text will be. Then those predictions magically turn into its outputs. See my earlier discussion: “LLM pretraining magically transmutes observations into behavior, in a way that is profoundly disanalogous to how brains work”. This leads to LLM behavior [...] --- Outline: (00:55) 1. Imitative learning → "seven deadly sins" misalignment (04:24) 2. Human approval → "glazing" misalignment (06:35) 3. Automatic verifiers → "literal genie" misalignment (08:05) 4. LLM judges → "trickster" misalignment (12:06) Afterword The original text contained 1 footnote which was omitted from this narration. --- First published: August 10th, 2026 Source: https://www.lesswrong.com/posts/GRmvZsHXH4vaijPMv/four-llm-loss-functions-four-flavors-of-llm-misalignment --- Narrated by TYPE III AUDIO.

Ratings & Reviews

4.7
out of 5
14 Ratings

About

Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.

You Might Also Like