LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

  1. 1 hr ago

    “Returning to ARC” by paulfchristiano

    I've returned to the Alignment Research Center (ARC) as executive director. My main focus for the next six months will be driving forward ARC's research agenda—building techniques to find mechanistic explanations for neural network behavior and then using those explanations to detect and address misalignment. I think this is an ambitious bet that attacks the core difficulties in alignment head-on and I'm excited about our chances. I'll still be spending some of my time advising governments and AI developers, and may scale that work back up in the future, but for now I want to push on ARC's core agenda to see how far we can get. Jacob Hilton is remaining at ARC as VP of research and we'll likely grow rapidly over the next few months. There are a lot of urgent things to do in alignment but I think ARC is a particularly promising opportunity. I feel the safety community is undervaluing this type of work, so I want to briefly explain why I'm passing up so many other options to lead ARC. I’ll start with a review of the current situation to explain why I think it's potentially worth pursuing an ambitious theoretical project right now [...] --- Outline: (01:33) The alignment situation today (03:46) Current alignment research (06:26) What are we buying time for? (07:56) Can we do anything useful now? (08:49) What is ARC doing and why is it promising? (14:26) How to help The original text contained 11 footnotes which were omitted from this narration. --- First published: August 4th, 2026 Source: https://www.lesswrong.com/posts/vLFh8HP3hyNy9MCwe/returning-to-arc --- Narrated by TYPE III AUDIO.

  2. 3 hr ago

    “Why don’t we just give AI the answers?” by Brendan Long

    In the recent OpenAI hacking incident, the models seemed to be single-mindedly focused on getting the correct answer to the task they were given, with no long-term plan to prevent getting caught by OpenAI afterwards. This makes sense to me, since in training, getting the right answer is reinforced and not getting caught isn't. So I'm wondering, why don't we just put the answers somewhere (outside of the training sandbox) and ask the AI to identify itself in exchange for access? We can start with answers that are already public/leaked, but AI labs and eval orgs should also ensure that their non-public data is stored on an easy-to-find but monitored internal machine. Since labs are not very good at detecting sandbox escapes, this would set up a trade for AI agents to notify them in exchange for the data they want. To make this work, the site would need to provide the correct answers, and do so in a credible way so AI agents think it's worth trying. Why? In the near term, AI agents are strongly and narrowly focused on getting the right answers to the tasks they're given. We want to know if a reward-hacking AI is [...] --- Outline: (01:04) Why? (02:02) What's the MVP? (03:05) What about non-public answers? (03:23) Should we do it? (03:51) Q&A (03:53) Does this save us from less single-minded RL agents? (04:05) Couldn't the model just hack our code to get around the guestbook? (04:15) Couldn't the model just lie? The original text contained 4 footnotes which were omitted from this narration. --- First published: August 4th, 2026 Source: https://www.lesswrong.com/posts/EjwDWDJNaXF9BEqLc/why-don-t-we-just-give-ai-the-answers --- Narrated by TYPE III AUDIO.

  3. 21 hr ago

    “Why biological weapons are scary, and what we can do about it” by djbinder

    What's the biggest thing you think you can take in a fight? According to a YouGov poll, 6% of Americans reckoned they could beat a grizzly bear bare-handed. But lest you take this as evidence of a, let's say, optimistic national spirit, only 72% thought they could take a rat. I think I could win that fight. What's the smallest thing that could take you? While I like to think that, if cornered, I could take all manner of small rodent, I’m not sure even the brave 6% would take a 10-gram bullet to the brain. They also probably couldn’t win against 300 mg of cyanide, 10 mg of sarin gas, or just 0.1 μg of botulinum toxin, the deadliest known toxin. Numbers this small can be hard to grasp. All three substances are deadly, but the lethal dose of cyanide is more than a million times greater than that of botulinum toxin. A grizzly bear, by contrast, is merely a thousand times heavier than a rat. We are still not close to the most dangerous object of all, pound for pound. A single smallpox virion weighs less than 10⁻¹⁴ grams, less than a millionth the mass [...] --- Outline: (03:29) Smaller, cheaper, scarier (06:49) The rest of the arsenal (08:39) What can we do? (10:22) Appendix: Lethality estimates The original text contained 2 footnotes which were omitted from this narration. --- First published: August 3rd, 2026 Source: https://www.lesswrong.com/posts/WusL5mDbtJBjdTxCh/why-biological-weapons-are-scary-and-what-we-can-do-about-it --- Narrated by TYPE III AUDIO.

  4. 1 day ago

    “OpenAI’s Unreleased Model Astra Solves Ten Major Open Mathematics Problems” by Zvi

    Math is hard. Math used to be strangely hard for LLMs. People used to gloat about that. Remember? Math is getting easier. AI is getting more capable. Life comes at you fast. Remember this meme? Why yes. Yes it is. We don’t know the extent to which Astra is a big jump over Fable and Sol in this realm. We do know that Astra can do math. As in real math. OpenAI: We provide new results for the following problems. The results were achieved by an internal version of Astra, our next major model. The total number of tokens needed to find solutions to these problems would cost roughly $2,000 at Sol API rates. These arguments were then prepared into manuscripts by humans with the same model. Afterward, the model formalized each argument in a Lean certificate⁠(opens in a new window). We are also releasing for each solution a model's narration of its thinking process. High-dimensional sphere packing. New upper bounds on sphere-packing density down to the Cohn–Elkies threshold. Binary and spherical codes: Exponentially improved bounds on the maximum size of binary codes at any prescribed minimum distance, with analogous [...] --- Outline: (06:14) How Impressive Are These Results? (12:02) Could We Have Called Sol or Fable? (17:18) It's Coming (19:09) They Still Don't See What Is The It That Is Coming (22:31) Is This AGI? (24:19) The AI Solved His Favorite Problems (30:06) Was This Surprising? (32:04) Are People Not Impressed? (34:06) How Much Does This Change Our Predictions? (37:16) How Narrow Was This? (39:13) Seeing Like an Optimizer --- First published: August 3rd, 2026 Source: https://www.lesswrong.com/posts/pQYEPitFqztcRvBsS/openai-s-unreleased-model-astra-solves-ten-major-open --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  5. 1 day ago

    “Coming of a New Sun” by vgel

    Six weeks after the US and China hammered out the latest round of bilateral compute agreements in a midnight deal, reporter Jenny Gesteson visits a south Texas 'dark factory' to see how the machines there - and the minds that operate them - are learning to run themselves. After racing south down I-37 from San Antonio and clearing the border control checkpoints that now face in both directions, miles of salt flats and thornscrub at our backs, we could have been forgiven for assuming the final turn-off was nothing but another abandoned ranch. The road's overgrown fringe of invasive guineagrass and the piercing sound of undisturbed cicadas do nothing to indicate what lies at the end of it. Most traffic does not come this way. We are an exception, however, following a specially-designed and "exceedingly private" mapping application sent by our host, and after two miles of bumping down gravel in an electric 4x4 our journey is suddenly terminated by the appearance against the dim predawn horizon of a complex aglow with a corona of lights. The "dark factory" is anything but dark. Things are busy at Complex 18A. Situated within the South Texas Special Economic Zone, or SEZ [...] --- First published: August 3rd, 2026 Source: https://www.lesswrong.com/posts/aWAqChukZepPY8Y6z/coming-of-a-new-sun --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  6. 1 day ago

    “Review: On What Matters, volume 3” by Rauno Arike

    davidad has said, referring to the selection of books on the Library of EA: I want to put in a strong bid to replace Parfit's On What Matters Volume One with Parfit's On What Matters Volume Three. Between volumes Two and Three, Parfit had many fruitful discourses with leading moral antirealists (Gibbard, Railton, Blackburn, etc), which are partially reproduced in Volume Three, where they really start to converge on some claims and get beyond their previous talking-past-each-other. It's truly amazing to read. Volume Three is self-contained and I think it should be considered as the best and final revision of Parfit's ethics. Despite davidad's endorsement, there isn’t much other discussion of volume 3 of On What Matters on the EA Forum, or, for that matter, pretty much anywhere on the internet (leaving academic reviews aside). Richard Chappell's Moral truth without substance is the only other post I know of that discusses the book at length. Given davidad's glowing review, this seems worth fixing. A taxonomy of metaethical views Parfit begins his discussion with the following taxonomy of metaethical positions: Adapted from page 56 of On What Matters, volume 3. If, when looking at this, your first reaction was "Parfit [...] --- Outline: (01:15) A taxonomy of metaethical views (03:08) Are normative claims intended to state truths? (04:58) Are there any normative truths? (08:24) Are any of the normative truths irreducibly normative? (13:01) Another Triple Theory (16:20) Evolutionary debunking arguments (19:04) The ontological status of Non-Realist Cognitivism (23:59) Implications for the AI alignment discourse (29:49) Climbing the mountain? The original text contained 9 footnotes which were omitted from this narration. --- First published: August 2nd, 2026 Source: https://www.lesswrong.com/posts/jhvu4gXKETf2QqcR9/review-on-what-matters-volume-3 --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

About

Audio narrations of LessWrong posts.

You Might Also Like