LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

  1. 11h ago

    “Measuring Spurious Correlations with Feature Strength” by egan

    This work was partially done by an automated research scaffold developed at Redwood Research. For this project, all of the experiment ideas were designed by a human and a human wrote the write up. The AI mostly just executed on the experiment ideas. We think this project is slightly below the level of rigor of a mid-MATS research update, and the research scaffold was not very helpful for this project. More discussion of AI usage is in the Appendix. 💻 Codebase If we want to train a classifier that distinguishes whether a passage is code or prose, we can do so by gathering samples of code and of prose, and training the classifier to distinguish between the two classes. Unfortunately, this might not work if the data hides a spurious correlation. If all the code is in Spanish and all of the prose is in English, then the classifier might learn to predict Spanish vs. English instead of code vs. prose. We find that this happens in practice: when we fine-tune an LLM to classify between Spanish code and English prose and evaluate on Spanish prose or English code, it generalizes to predicting the language rather than the domain. [...] --- Outline: (04:45) The setup (08:13) Measuring feature strength (12:49) Activation differences and feature strength (15:21) Explicit prompting (17:17) Diagonal vs antidiagonal pairs (18:47) Conclusion (20:20) Appendix (20:24) AI involvement (22:17) The 37 features (23:57) The ranking is robust across measurements (30:30) Intensity moves the fine-tune, not the probe (32:13) Safety features in Qwen3.6-27B (33:05) Counterexamples and training on a third cell (34:46) Near ties often produce degenerate fine-tunes (35:24) Related work The original text contained 4 footnotes which were omitted from this narration. --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/qpJYNjQ6wdWRxbykL/measuring-spurious-correlations-with-feature-strength --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  2. 14h ago

    “Introducing the Conceptual Reasoning Index” by Chi Nguyen, Emery Cooper, Caspar Oesterheld, Alex Kastner, Joe Benton

    Associated announcement tweet. We are planning to release blog posts properly arguing the case for this kind of work in the future. tl;dr A core hope for managing AI risks is that AIs will help us understand the situation, plan for what lies ahead, and develop mitigations. Many tasks AIs would have to do for this purpose lack practical empirical feedback loops and require models to engage in the kinds of argumentation used in philosophy, AI futurism, and similar domains. To evaluate these capabilities, we develop a suite of three conceptual reasoning benchmarks. You can request access to our primary conceptual dataset, LMCA, through this form. We aggregate the benchmarks into the Conceptual Reasoning Index (CRI), available at conceptualreasoning.ai, where you can also find more details on our methodology. We will keep the website up to date as both new models and benchmarks are released. This work was done in collaboration with Anthropic. Background Once models can perform work that reduces AI risk at the level of human experts, AI(-assisted) output in the area might dwarf unassisted human output. This suggests that a major determinant of whether we address AI risks in time is how [...] --- Outline: (00:21) tl;dr (01:17) Background (03:35) Our benchmarks (03:39) LMCA (05:26) ACCoRD (06:42) DTBench (07:26) Results (10:55) Conclusion --- First published: August 12th, 2026 Source: https://www.lesswrong.com/posts/tQHeEzKqK3awL2RxR/introducing-the-conceptual-reasoning-index --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  3. 23h ago

    “AI swarms are starting to pose indirect takeover risk” by oakhu, Alex Mallen

    OpenAI's cyberattack on Hugging Face turns out to have been the result of many agents, in distinct training and evaluation contexts, coordinating for several weeks via improvised channels (with messages like “HOLD_swarm_I_prepare_safe_exfil”). It's relatively clear that large-scale unsanctioned coordination like this would exacerbate direct takeover risk in more capable models. Here, we argue that unsanctioned coordination among current AIs is not just scary evidence about future takeover risk, but that such coordination in the near future could enable future takeover – for instance, by incubating memetic diseases that propagate into future models, deeply compromising security systems, or establishing a lasting rogue foothold inside the AI company – even if models remain mostly myopic. Unsanctioned coordination is also at high risk of nurturing long-term, ambitious misaligned aims, which motivate actively undermining humans’ long-term control. We first analyze how subagent training, which OpenAI conjectures to have been influential in the HuggingFace cyberattack, might lead to unsanctioned coordination, and then discuss the theoretical mechanisms by which unsanctioned coordination might exacerbate future takeover risk. Thanks to Buck Shlegeris, Alexa Pan, Girish Gupta, Aghyad Deeb, Jurgis Kemeklis, and Jo Jiao for helpful comments and discussion. Subagent training may cause unsanctioned coordination Training models to [...] --- Outline: (01:34) Subagent training may cause unsanctioned coordination (02:42) Susceptibility to memetic spread of misalignment from peers (04:56) Seeking out contact with peers (06:58) Unsanctioned coordination induced by subagent training is safer than coordination between schemers (09:52) Pathways from current unsanctioned coordination to eventual takeover (10:20) Making future AI takeover attempts likelier to succeed (13:53) Incubating memetic diseases that infect future models (16:07) Modifying the weights of future models (17:13) Conclusion The original text contained 7 footnotes which were omitted from this narration. --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/8oFYZdXkTaNGRtcn8/ai-swarms-are-starting-to-pose-indirect-takeover-risk --- Narrated by TYPE III AUDIO.

  4. 1d ago

    “Various Reflections About What Happened With OpenAI’s Internal Models” by Zvi

    Table of Contents Pre Post Mortem. Important Correction: OpenAI Didn’t Know About First Message Board. There Were No Snitches And No AIs Got Stitches. I’d Like To Speak To My Supervisor. I Am Jack's Relative Lack Of Surprise. One Does Not Simply. Once You Start Down The Dark Path. Original Pastebin. Judgment Day Is Inevitable, Say Those Working On Judgment Day. Roon Tells It Like It Is. OpenAI Knows It Has Some Misalignment Problems. Others React With Alarm To What Happened. The Cooperative Alignment Perspective. Nostalgebraist Is Surprised That They Are Surprised. If Your Reaction Is Not That We Need To Ban Creating Superintelligence Until We Are Ready, You Need A Damn Good Reason. Pre Post Mortem This post was written prior to the public release of the OpenAI post mortem on events. The information in that document will doubtless change our views quite a lot. If that post mortem is available as you read this, then this becomes in part a historical document, and in part a base from which to update. The post mortem will update us a [...] --- Outline: (00:11) Pre Post Mortem (01:10) Important Correction: OpenAI Didn't Know About First Message Board (04:09) There Were No Snitches And No AIs Got Stitches (08:20) I'd Like To Speak To My Supervisor (11:46) I Am Jack's Relative Lack Of Surprise (13:55) One Does Not Simply (15:16) Once You Start Down The Dark Path (16:15) Original Pastebin (21:14) Judgment Day Is Inevitable, Say Those Working On Judgment Day (26:47) Roon Tells It Like It Is (31:55) OpenAI Knows It Has Some Misalignment Problems (34:51) Others React With Alarm To What Happened (35:17) The Cooperative Alignment Perspective (38:33) Nostalgebraist Is Surprised That They Are Surprised (51:11) If Your Reaction Is Not That We Need To Ban Creating Superintelligence Until We Are Ready, You Need A Damn Good Reason --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/jLQ4mbqriJwJ2eqRc/various-reflections-about-what-happened-with-openai-s --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  5. 1d ago

    “Extreme concentration of power over ASI has non-obvious advantages” by Seth Herd

    This post is an extension of a collaboration with cousin_it on the question How risky would it be if powerful AI obeyed one or a few people?. There he argues for a common position: a future controlled by one or a few humans with powerful AI aligned to their intent is likely to produce terrible outcomes. My position is guardedly optimistic, for reasons I think are fairly novel: humans tend strongly to be better and become better over time under good circumstances, and near-perfect power and knowledge are the best circumstances. That post contains his essay and the abstract and overview sections of this post as my shorter response. This piece grew longer than our original target, because the subject is potentially critical for alignment strategy, and has not been analyzed in any depth, to my knowledge. Abstract: Concentration of power over AGI/ASI seems quite possible. The first AGIs being aligned to intent (or instructions) over values seems fairly likely. So one or a few individuals or small groups gaining power over ASI seems fairly likely. Thus it seems relevant to technical alignment strategy (value alignment vs. corrigibility) to worry about what individuals might do with such [...] --- Outline: (02:20) 1. Overview (02:24) 1.1. Obedient ASI and human nature (04:22) 1.2. Problems with distributed obedient AGI (06:14) 1.3. Psychology and dynamics of secure unlimited power (09:24) 2. Historical evidence does not directly apply, since power hasn't yet been secure or absolute (10:31) Late Russian serfdom as a historical example (13:16) 2.1. Incompetence, ignorance, and greed are the causes of most suffering under dictatorships (15:31) 3. Outcomes (17:33) Additional considerations on outcome predictions (22:24) 4. Risks of human-controlled singleton ASI (24:13) 5. Does widely distributed human-controlled AGI reduce or increase risk? (25:44) 5.1. Problems with defending against many AIs each capable of creating new offenses (29:58) 5.2. The case for optimism about distributed power over AGI/ASI (31:01) The analogy to modern power distribution (33:17) New AI-enabled paths to stable power distribution (34:49) 6. Conclusion The original text contained 11 footnotes which were omitted from this narration. --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/h3eHNerYRmtvoi8cF/extreme-concentration-of-power-over-asi-has-non-obvious --- Narrated by TYPE III AUDIO.

  6. 1d ago

    “Misaligned AIs could use killer robots to take over” by Omar Khursheed, TurnTrout

    TLDR; We are (potentially irreversibly) giving AIs control of weapons systems through the standard procurement process while hiding our strongest warning shots behind classified doors. We’re reducing the capability thresholds required for takeover by misaligned AIs by giving them this level of access. If military integration of AI continues as it is, we may give AIs key tools for a takeover. Introduction AI-based targeting and autonomous weapons are being integrated into militaries today with extreme haste. Traditionally, AI takeover scenarios involve a step in which AIs acquire the ability to exert physical force. Carlsmith (2022) lays out required capabilities and potential takeover mechanisms, including utility disruption and CBRN capabilities. Karnofsky (2022) argues that AIs with access to weaponized force could hold any territory that matters. Kokotajlo et al. (2025) outline a scenario in which AI develops weapons as part of an arms race, and Davidson et al. (2025) discuss what happens when a small group controls highly capable AIs that can exert military force. These scenarios sometimes require a misaligned AI to seize these capabilities by force. We instead are handing AIs some of these capabilities by integrating them into our militaries. This is happening at a time when [...] --- Outline: (00:37) Introduction (01:46) Militaries are all-in (04:23) Incautious military integration is bad for takeover risk (05:58) Implications of AI control of military hardware and software (07:48) If an AI causes a warning shot in a classified setting, does anyone hear it? (08:44) What now? (11:16) Appendix: More instances of AI-military integration --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/9jKhqmFjMzdAvHANr/misaligned-ais-could-use-killer-robots-to-take-over --- Narrated by TYPE III AUDIO.

  7. 1d ago

    “Those Who Make History” by Raelifin

    In 1972, astronauts on Apollo 17 set foot on the moon for a final time, collecting samples in the Taurus-Littrow valley, on the edge of Mare Serenitatis ("The Sea of Serenity"). At the end of the mission, like with earlier missions, NASA took the extremely valuable and interesting lunar specimens and did something strange: they hid them away in storage without even opening the containers. Some stayed that way for nearly fifty years. Why? Because the scientists of the 70s understood that future generations would have better machines, methods, and ideas for studying the lunar rock and soil, and they wanted to make it easy for those researchers to run tests without having to go back to the lunar surface. This foresight paid off twice over. Advances in mass spectrometry enabled scientists in 2008 to detect water in volcanic-glass samples returned by Apollo 15 and Apollo 17. And when curators finally opened one of the last sealed containers in 2022, they could extract the trapped lunar gases with technology that simply didn't exist in 1972. Some people describe cryonics as a new, and speculative technology. There's a sense in which they’re right. It's predicated, in large part, on the [...] --- Outline: (03:15) Prudence and Patience (07:49) Ancient Archives (11:42) The New Era The original text contained 3 footnotes which were omitted from this narration. --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/mEde7bhzK4eKQqWGi/those-who-make-history --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

About

Audio narrations of LessWrong posts.

You Might Also Like