LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

  1. 5h ago

    “The Game is Set for a Targeted Memetic Attack on the AI Safety Community” by keltan

    While this is relevant to my work at MIRI, I have not checked these ideas with anyone else on the team and am posting this on my personal LW account. These views are my own. And to be honest, I am writing this mostly to remind myself of my weakness. --- I expect one (or many) adversarial memetic attacks aiming to trip you up, perhaps consisting of fake leaks relating to dangerous stuff happening in the labs. Specifically, worrying incidents that may fit snugly within your worldview, leaking from multiple sources including news outlet/s, but not confirmed/confirmable by a primary source. Think rumors about exfiltrated weights, AIs attempting to create viruses, agent swarms hacking into and gathering information from nuclear infrastructure, etc. An easy way to remove status from a movement is to trip it up: make it fall for a misinformation trap in public, then use that slip-up to discredit the movement for all time. The game is set for a memetic attack like this. There's a well-resourced group waiting for your screw-up. And then you may remember much that will help you.  In public and in private, if you feel surprised or confused, notice your confusion. These [...] The original text contained 2 footnotes which were omitted from this narration. --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/5mcDjo5gjn3Leahhu/the-game-is-set-for-a-targeted-memetic-attack-on-the-ai --- Narrated by TYPE III AUDIO.

  2. 9h ago

    “Deep recurrent models are less robustly CoT-monitorable than normal CoT models in a toy setting” by Nick Kuhn, Alek Westover

    We use RL to teach a deep recurrent model and a normal CoT model to solve a math problem while hiding from a CoT monitor which of two possible problems it is solving. The deep recurrent model moves its reasoning into latents (thereby evading the CoT monitor) within 40 steps of RL; the CoT control is much worse at learning to confuse the LLM monitor. Note: in the above graph, “monitorability” is how much better than chance the CoT monitor is at figuring out which problem the model is solving. This was our simplest setting; similar settings show similar results with somewhat smaller differences between the CoT model and the deep recurrent model, see Results. Some proposed architectural changes add a latent channel in addition to the normal text channel for persisting state between forward passes. These parallel latents architectures (like T2MLR and full-bandwidth transformers) could theoretically perform all their reasoning in latents, rather than text. We’ve speculated that these architectures will therefore be harder to oversee. This post empirically investigates whether these models can learn to reason without their CoT (which would likely make them harder to oversee) in a simplified setting. We compare the performance of a [...] --- Outline: (03:49) A formal description of the parallel latents architecture (05:00) Experiment 1 (05:50) Tasks (09:09) Stage 1: SFT (11:53) Stage 2: RL (13:06) Results (20:50) Discussion (21:55) Task (22:50) Pressure to solve the task (23:32) Oversight method (24:25) Model (25:53) Obfuscation pressure (28:04) Suggestions for future work (29:58) Conclusion (30:26) Acknowledgements (30:49) Appendix (30:52) Experiment 2: Output Spillover (32:22) Result (33:25) Experiment 3: Training pressure variations (37:09) en-US-AvaMultilingualNeural__ When starting with low task performance, rejection filtering on highly monitorable trajectories can lead to increased obfuscation. (37:20) en-US-AvaMultilingualNeural__ When initial performance on the task is high, rejection filtering does not exert significant pressure on monitorability. The original text contained 3 footnotes which were omitted from this narration. --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/5guQJSqstkjgys3PE/deep-recurrent-models-are-less-robustly-cot-monitorable-than --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  3. 11h ago

    “Superintelligence this Christmas” by Alexander Gietelink Oldenziel

    I now consider it plausible that some form of recursive self-improvement is imminent, and that we may be on track for superintelligence by Christmas of this year if racing continues. This is substantially faster than any forecast, including ones like AI 2027 that were considered outrageously fast a year ago. It is faster than I myself expected even a week ago. I don't work at a scaling lab. I don't know more than is public knowledge. Let me be perfectly clear: what I am saying is absolutely nuts. Extraordinary claims require extraordinary evidence. I claim we have now received said evidence and you should update accordingly. FOOM should probably should be your *default expectation*. People have strong status quo bias. Your default expectation should be that things will radically speed up. We are not at the ceiling of intelligence. We should probably expect the transition to superintelligence to be incredibly fast. RSI is a positive feedback loop, so it is inherently (hyper)exponential. Everything is an S-curve eventually, but nothing suggests the ceiling is anywhere near human level, or that it happens at a human timescale. AI is [...] --- Outline: (01:08) FOOM should probably should be your *default expectation*. (01:44) AI is capable of revolutionary advances in mathematics. Machine learning research is not different in kind. (03:16) The speed of AI progress continues to be underestimated; by superforecasters and even by the researchers themselves. (05:07) Internal models are significantly ahead of released ones; (06:08) Intuitions about timing from pre-training runs are misleading since most progress comes from RL, unhobbling and algorithmic innovations (06:29) Enter the Swarm (07:04) Anthropic's own report states it has 30,000 agents running concurrently, and Claude has completely taken over 26% of all R&D. The original text contained 3 footnotes which were omitted from this narration. --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/LJbKwctaioqp2Hi4b/superintelligence-this-christmas --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  4. 14h ago

    “Grantmakers aren’t afraid to die” by dan.parshall

    The AI Risk grantmakers do not act like they believe in imminent existential risk from AI The idea of "revealed preferences" is one of the most useful in economics; it allows us to cut through a great deal of metaphysical angst about what someone "really" believes, and focus on what they act like they believe, which is much more useful for making predictions about their future actions. As one example, I grew up in a, shall we say, fervently-religious community, and it's often hard for nerdy Rationalist types to understand this, but: there are people who genuinely believe in Hell, and in Heaven. They genuinely believe that moving souls from one to the other is the most important thing on Earth. It's one thing to doubt the conviction of someone who lives an easy, staid, middle-class life... but for others, their choices and behaviors (e.g. years-long missionary trips) reveal their true preference and/or belief beyond any reasonable doubt. I bring this up because, per the actions and decisions of grantmakers operating in the AI Risk space, they mostly DO NOT seem to believe in imminent existential risk of AI. On the contrary, they act like people who [...] --- Outline: (00:10) The AI Risk grantmakers do not act like they believe in imminent existential risk from AI (01:27) The explore-exploit tradeoff (02:44) The evidence we're in "exploit" mode (02:49) Exhibit A (03:09) Exhibit B (03:32) Exhibit C (04:32) Obvious verdict is obvious (06:35) Explore mode: Just do (good) things (better) (09:35) Conclusion The original text contained 13 footnotes which were omitted from this narration. --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/2o8B9hDN94k4Qine6/grantmakers-aren-t-afraid-to-die --- Narrated by TYPE III AUDIO.

  5. 14h ago

    [Linkpost] “Pacing the Frontier: A Framework & Research Agenda” by CharlesD, technicalities, Raymond Douglas, Nowe Moore

    This is a link post. Below is the executive summary from our new paper at pacing.tech. The full paper is available on the site and as a PDF. The full author list is Raymond Douglas, Charles Dillon, Nikola Moore, Gavin Leech, Shahar Avin, Mathias Kirk Bonde, Rohit Krishnan, Noah Perez, Nathan Young, Cormac Slade Byrd, Stephen Casper, Jan Kulveit, & David Duvenaud “Pacing AI” usually refers to how to conclusively handle the most extreme risks in the face of race dynamics. However, even for the goal of handling these highest-stakes cases, it's useful to take a broad view of pacing—one that encompasses all interventions aimed at moderating the pace of AI development, deployment, or diffusion. Thus: Haphazard pacing is already common, including: delaying model releases for safety testing, pausing model development in response to shocks, and applying export controls.Current approaches will predictably fail. Isolated, unilateral actions addressing only small fractions of the problem are not enough, but poorly executed interventions could easily backfire—good solutions will need to be carefully designed.Precedents are being set whether we like it or not. How AI progress is paced now will shape how it is paced in future. We can learn [...] --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/E5SmpFsGPNpYjf92c/pacing-the-frontier-a-framework-and-research-agenda Linkpost URL:pacing.tech --- Narrated by TYPE III AUDIO.

  6. 17h ago

    “plzdontkillus Fellows Got ~2M AI Safety Views, Not 21M” by Josh Thorsteinson

    Summary I was a fellow at plzdontkillus, a month-long creator bootcamp at Lighthaven, partially funded by MIRI, where ~55 fellows posted one video per day. plzdontkillus.com originally claimed “21M+ AI risk views” with no breakdown. After I shared a draft of this post, the organizers relabeled it “X-Risk Relevant Views” and published one. Three videos account for 80% of the views: a datacenter-water-use debunk (8.5M), an AI dystopia video (6.4M), and a Rob Miles Hugging Face incident explainer (2.5M). The rest total 4.3M. Under my stricter definition of AI safety content, fellows generated ~2 million views total. Based on my analysis, fellow-made AI safety videos made up around ¼ of fellows’ output and ~2% of total views. 13 out of ~55 fellows posted zero AI safety videos, and an additional 8 posted only one or two. This is partly because the program didn't incentivize AI safety content. If they run it again, I think they should change that. Me I’m Josh Thor. I was a fellowLike every fellow, plzdontkillus offered me a $2000 stipend and free room and board for the month (which I accepted)I won the program's “Other” category for [...] --- Outline: (00:13) Summary (01:28) Me (02:12) What they claim (05:26) My analysis (06:58) Program incentives (08:38) Aella's response (11:00) My recommendation The original text contained 11 footnotes which were omitted from this narration. --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/LQ9wKT9oNeArbwukz/plzdontkillus-fellows-got-2m-ai-safety-views-not-21m --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  7. 17h ago

    “Obstacles to the scalable oversight of auto-alignment research” by Sam Martin, Dewi Gould, Cameron Holmes, Jacob Pfau

    TL;DR. In this work we study obstacles to the faithful automation of alignment research. We see this as a scalable oversight problem. There are plenty of examples of how models fail at this, and as models become more capable our ability to notice these failures will diminish: even the best human checkers won’t be able to tell if the model was well elicited, thorough checking will become too costly, and models could tailor their responses to their judges. We draw on empirical examples from Geoguessr and auto-alignment runs from Arcadia's internal research to make general claims about obstacles to the oversight of fuzzy alignment-related tasks. Narrowing our attention to one prominent scalable oversight method, we find that whilst debate shows promise on typical capabilities benchmarks (aligning with recent work) it fails on tasks involving judgment calls akin to those arising in automated alignment research. We’d like to thank David Africa, Andrew Draganov, Rory Greig, Joshua Jacob, Rishub Jain, Zac Kenton, Francis Rhys Ward and Lennie Wells for helpful feedback on this post. Introduction Existing empirical work on debate [1, 2, 3, 4, 5, 6] has almost exclusively focused on objective, verifiable domains, seeking to mitigate misalignment caused by supervision [...] --- Outline: (01:20) Introduction (04:00) Decomposition of explanations (10:00) Empirical Examples (10:18) Geoguessr Setting (11:59) Example claims in fuzzy arguments (12:29) Nature of arguments in non-fuzzy tasks (14:39) Discussion: scalable oversight of fuzzy tasks (17:05) Empirical Debate Results (17:09) Geoguessr (18:39) LMCA Debate (19:43) Conclusion (20:18) Appendix (20:21) Geoguessr Setting The original text contained 7 footnotes which were omitted from this narration. --- First published: September 16th, 2026 Source: https://www.lesswrong.com/posts/PBGKWNrJAbpDgSsPo/obstacles-to-the-scalable-oversight-of-auto-alignment --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

About

Audio narrations of LessWrong posts.

You Might Also Like