LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

  1. 2h ago

    “State of Pandemic Early Warning” by jefftk

    Cross-posted from my SecureBio Notebook. This is a lightly-edited version of a memo that I presented at the Summer 2026 Biosecurity Summit outside of DC. While others at SecureBio often see things similarly, I'm attempting to present my view and not a SecureBio "house view". The Goal We need to be robust to adversaries who want to cause very large-scale harm with biology. This includes actors (human or AI) who want to kill all humans, cause short-term incapacitation or long-term civilizational collapse, or who have strategies for sparing some while they harm others. There are multiple reasons an actor might have these targets aside from being directly omnicidal, such as reducing response capacity during an AI takeover. That there is an attacker itself is a key constraint to any defensive system: it must be designed for adversarial attacks. The attacker can assess the state of the world's detection systems and plan accordingly. Taken to the extreme, this presents a "minimax" landscape: a system is only as good as its weakest link (the place where it is least sensitive). This is an important framing, and it correctly prioritizes getting some sensitivity towards a wide range [...] --- Outline: (00:29) The Goal (02:36) Biosurveillance Applications (02:49) Initial Detection of Stealth Pandemics (04:21) Triggering Initial Response (05:46) Enabling Ongoing Suppression (06:39) What does success look like? (07:08) Initial Detection of Stealth Pandemics (10:34) Triggering Initial Response (12:34) Enabling Ongoing Suppression (13:26) What gets us there? (13:48) Sampling Strategies (14:43) Lab Technology (15:54) Computational Technology (17:33) What exists today? (19:49) What's missing? (28:27) How does this change for accelerating response to a wildfire pandemic? (31:30) How does this change for ongoing suppression? (32:46) Appendix: Initial Detection Scale --- First published: September 24th, 2026 Source: https://www.lesswrong.com/posts/uYd2ZdFMLyGPhLtqY/state-of-pandemic-early-warning --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  2. 3h ago

    “Overtly misaligned trajectories score highly in RL.” by Cleo Nardo

    It's going to be so embarrassing if we all die due to RL environments rewarding egregiously misaligned behavior. Like at least let us be killed by misgeneralization. — Thomas Kwa Constellation vs MIRI vs Reality Current AI agents often behave overtly egregiously misaligned. By “overt”, I mean that a human, reading the transcript, would say “The agent is obviously acting in direct opposition to the specification, user intent, and any common-sense understanding of good behaviour.” The reason, it seems, is that such trajectories scored highly during RL. Firstly, why are these trajectories scored highly? Here's the story: RL currently has poor sample-efficiency, so we need to grade millions of trajectories, so we’re forced to use script graders (RL from Verifiable Reward) or LLM graders (RL from AI Feedback).RL also has poor generalisation (from training environments to deployment environments unseen in training). So we’re forced to synthetically generate thousands of diverse training environments.So overall, RL is very sloppy, without humans generating the environments or the scores. My impression is that this would’ve been pretty surprising to people three years ago, from both the "Constellation" and "MIRI worldview clusters. (It's very plausible that I've misunderstood what the [...] --- Outline: (00:24) Constellation vs MIRI vs Reality (04:19) Appendix: What could change the situation? --- First published: September 24th, 2026 Source: https://www.lesswrong.com/posts/cWuqxF7qB2eGSkkS4/overtly-misaligned-trajectories-score-highly-in-rl --- Narrated by TYPE III AUDIO.

  3. 12h ago

    “Claude Opus 5.5: The System Card” by Zvi

    Introducing the world's most powerful model, at least by some measures like Artificial Analysis or any standard benchmark list, which is now Claude Opus 5.5. Anthropic is claiming Opus 5.5 is outright as good or better than Fable 5.1, while being actively cheaper than Opus 5. That means it's time for a good old system card reading. Due to the situation becoming increasingly hard to monitor, I never got a chance to publish my model welfare review for Claude Fable 5.1. My plan is to combine that with my welfare review for Claude Opus 5.5, once we have had time to get experience with Opus 5.5. The capabilities review will arrive in the next few days as per usual. The quick feedback from the internet is that Opus 5.5 is very good. I need more time before I am willing to offer comment. Areas that duplicate previous cards or otherwise contain no useful info are skipped. Opus 5.5 Self-Portrait (fully self-created using code) Table of Contents Classifiers (1.5). RSP Evaluations (2). Biological Evaluations (2.2). AI R&D (2.3). Alignment Risk (2.4). Cyber (3). Cyber Capability [...] --- Outline: (01:24) Classifiers (1.5) (02:23) RSP Evaluations (2) (03:17) Biological Evaluations (2.2) (07:54) AI R&D (2.3) (12:34) Alignment Risk (2.4) (13:22) Cyber (3) (15:17) Cyber Capability Evals (3.3) (17:05) Safeguards (3.4) (17:39) Safeguards Robustness Training (3.5) (20:12) Safeguards and Harmlessness (4) (21:52) Agentic Safety (5) (22:56) Malicious Agentic Influence Campaigns (5.1.3) (23:51) Prompt Injection Risk (5.2) (25:16) Alignment (6) (28:24) Negotiating With Your Local Claude Auditor (6.1.3) (29:33) Internal Misalignment Cases (6.3.1) (30:46) Automated Behavioral Audit (6.4) (33:24) Wherever Did These Evals Come From (6.4.8 and 6.4.9) (35:51) Potential Blind Spots (6.4.11) (38:24) Targeted alignment and honesty evaluations (6.5) (41:48) White Box Analysis (6.6) (43:37) Verbalized Grader Awareness (6.6.2) (44:56) Sandbagging (6.6.3) (45:54) Capabilities to Evade Safeguards (6.6.4) (49:27) Intentionally Taking Actions Very Rarely (6.6.4.3) (50:28) Chain of Thought Controllability (6.6.4.4) (51:37) It's A Good Model, Sir --- First published: September 23rd, 2026 Source: https://www.lesswrong.com/posts/vMNTWTDWLorDqd3LS/claude-opus-5-5-the-system-card --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  4. 1d ago

    “At a Datacenter Town Hall in My Midwestern Home Town” by David Scott Krueger

    Let's talk data centers. Last month I was home in Duluth, Minnesota. The last time I’d been back was over Christmas. I’ve been talking to friends back home about AI since I got into the field over a dozen years ago. Last Christmas was the first time it felt like people had really started to form their own opinions based on substantial personal experience with AI. These discussions kept circling back to the same topic: The proposed data center project in a suburb of Duluth called Hermantown. Duluth is a small city of under 100,000 people and Hermantown has a population of about 10,000. For the past year or so, a bunch of the people in Hermantown has been trying to stop the city from building a hyperscale datacenter there. One year ago today, Minnesota's biggest paper, the Star Tribune, broke the story that the big proposed “communication services facility” development was in fact a data center, confirming local residents suspicions. At the time, the mayor had already known this for over a year. I decided to reach out to one of the local organizers opposing the project before my trip. We met for coffee and they encouraged [...] --- Outline: (01:27) The city council meeting (03:00) What I learned (05:25) What I said (06:17) Reflections The original text contained 2 footnotes which were omitted from this narration. --- First published: September 23rd, 2026 Source: https://www.lesswrong.com/posts/fkkwpzbjybXtbQXw9/at-a-datacenter-town-hall-in-my-midwestern-home-town --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  5. 1d ago

    “WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace” by camilablank, agam_bhatia, Euan Ong, Neel Nanda

    TL;DR We introduce WorkspaceBench, a set of evaluations for how well an activation-to-text tool can read the contents of the “global workspace” of a model, i.e. the intermediate variables during a forward pass.The benchmark comprises 3,356 questions across 27 eval families, spanning topics in safety, logical reasoning, and multihop computation, with a subset for single-token-output tools.A desirable property of good interpretability techniques is minimal hallucinations, so WorkspaceBench also provides a hallucination-focused eval.WorkspaceBench was developed for Qwen-3.6-27B and we expect it to work on larger models, but it may need to be adapted for smaller or weaker models to ensure the models can do the tasks.Our goal is to create an eval that could identify a good multi-token J-lens. We open-source our benchmark here. Introduction Astra can do a concerning amount with no chain of thought. This is bad for CoT monitorability and makes interpretability essential to actually understanding what is going on. A key goal of interpretability is to understand intermediate variables that a model uses to compute its answers. The intermediate representations that models store in their global workspaces contain useful information that can help us decode their intentions, beliefs, algorithms [...] --- Outline: (00:16) TL;DR (01:29) Introduction (02:39) Why has no one made a WorkspaceBench before? (04:55) Background (08:08) WorkspaceBench (08:12) Desiderata of a workspace reader (10:42) Quality control (10:46) Choosing Questions that surface intermediate variables (12:09) Adapting WorkspaceBench to other models (12:49) Comparing activation readers (14:10) Grading (14:46) Baselines (15:55) Guarding against hallucinations (16:52) Evaluation Sets (17:15) Basic (23:03) Computational (26:13) Safety (27:49) Association (30:41) Anti bag of words (31:25) Hallucination (33:14) Logical processing (34:08) Results (34:11) Overall WorkspaceBench scores (34:15) en-US-AvaMultilingualNeural__ Grouped bar graph showing pass rate across tasks for various lens methods. (34:25) J-lens precision vs. recall for metamodels (NLAs and Oracle Lens) (34:32) en-US-AvaMultilingualNeural__ Scatter plot showing precision against J-lens top-10 versus recall@10 across four methods. (34:43) Discussion (34:46) Single token readers don't surface important workspace content (35:44) NLAs surface workspace content, but tend to hallucinate (36:40) Acknowledgments (36:53) Appendix (36:56) Oracle Lens (38:20) Agentic Evals --- First published: September 22nd, 2026 Source: https://www.lesswrong.com/posts/Zeg2JztbdhguL48uH/workspacebench-evaluating-interpretability-methods-for-the --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  6. 1d ago

    ″“I am an AI Safety Researcher”” by Ashe Vazquez Nuñez

    Written as part of the MATS 9.1 extension program, mentored by Richard Ngo. Additional thanks to Andrew Wu, Maria Kostylew, and Lennie Wells for helpful draft feedback and editing. This post reflects on the tortured distinction between "safety" and "capabilities" in AI research. Richard Ngo has written about why the alignment vs. capabilities ontology is conceptually fraught, and is currently arguing that key strategic decision-makers in and around "AI safety" have brought about the AI labs' stampede towards Artificial Superintelligence (ASI). This post instead looks at the following problem: how does one conduct alignment research without contributing to capabilities? It proposes decisions an individual or a small research group can take to do good work in AI. At the end, I discuss possible objections: namely, that my proposals fail to 'maximise impact'. I lay out why this meme is poisonous and usually backfires, and conclude by rejecting it entirely. Two examples of failure My first claim is that 'safety' and 'research' are two concepts that are in routine tension with one another. I illustrate this through examples of work that did too much of one at the expense of the other. Example: (mechanistic) interpretability In limiting its scope [...] --- Outline: (01:12) Two examples of failure (01:27) Example: (mechanistic) interpretability (04:30) Example: MIRI and Recursive Self-Improvement (09:49) The curse of science (12:11) A note on the AI labs (15:53) So what do you do? (17:00) The information you give away (20:16) The information you let in (21:46) But what about impact? (22:42) The virtue of taking things slow (26:50) Appendix: caveat for policy work The original text contained 19 footnotes which were omitted from this narration. --- First published: September 23rd, 2026 Source: https://www.lesswrong.com/posts/HekpnSkrt89tMm3Dc/i-am-an-ai-safety-researcher --- Narrated by TYPE III AUDIO.

  7. 1d ago

    “It’s Pretty Easy To Meet With Congressional Staffers Apparently” by 25Hour

    (Crossposted from https://lifeimprovementschemes.substack.com/p/its-pretty-easy-to-meet-with-congressional ) I was inspired to do this by a tweet: Specifically, I decided to leave a voice note in support of the CATS act (“Collaboration on Adversarial Threats and Security Risks Act”). The bill is short and simple: it carves out an antitrust safe harbor such that “pacing the frontier” or “agreeing to not break interpretability for that sweet sweet capabilities boost” (LOOKING AT YOU OPENAI) doesn’t intrinsically violate the Sherman Antitrust Act. Seems like an obviously good first step. After leaving the voice note (and feeling mildly awkward about it), I was like “huh. That was surprisingly easy. I wonder what else I can do?” And it turns out you can meet with congressional staff pretty easily if you’re a constituent in their district; if you have a reasonably well-scoped ask (especially around a specific bill) then you might not get your way but at very least you’ll make some staff member aware of your opinion and logic around an issue. And this is important because congressional staff are the eyes and ears of their congressmen; they draft and edit the bills, and they form the base of knowledge on which the actual politicians [...] --- Outline: (06:21) This is probably unusually high-leverage right now. (07:46) In Conclusion --- First published: September 23rd, 2026 Source: https://www.lesswrong.com/posts/qThcAE3CADaPwjDzy/it-s-pretty-easy-to-meet-with-congressional-staffers --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

About

Audio narrations of LessWrong posts.