LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

  1. 1h ago

    “Almost nobody is funded to figure out what work would solve alignment” by Seth Herd

    Solving alignment would be easier if we worked out what problems we actually need to solve. This could be called the alignment meta-problem. Work on this problem is rarely directly funded. More focused work on it should let us use our limited time and funding more efficiently. The diagram implies narrowing alignment work, but I expect meta-problem work to also identify high-payoff "fringe" approaches. If we're driving toward a cliff, maybe we should buy better headlights. All too often we're doing work that merely sounds or feels good, and optimizing less than we could for work that drives most efficiently toward success. Some of this is inevitable and some of it is useful, but we could do more to light the path ahead. Most researchers agree that mech interp, refining and improving alignment training, control, theory, and miscellaneous techniques like confession are useful for solving alignment. Working toward regulation and slowdown/pause is also commonly considered useful in the governance space, and spreading awareness of alignment risks is pretty obviously useful for accelerating and enabling all of this work. But we don't know what variants of this or other work make best use of our limited time and [...] --- Outline: (03:03) Why not to fund more work on the meta-problem (03:50) Arguments in favor, compressed --- First published: September 4th, 2026 Source: https://www.lesswrong.com/posts/g4eaRCynouiBi2LjQ/almost-nobody-is-funded-to-figure-out-what-work-would-solve --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  2. 2h ago

    “AI #184: Post Post Mortem” by Zvi

    I am exhausted. We may finally be nearing the end of direct coverage of What Happened with the attack on HuggingFace, and the subsequent near term reactions. That took up a full five posts in the last week: OpenAI Offers Straight-Laced Postmortem of the HuggingFace Hack. METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack. HuggingFace Attack Postmortem: Fleshing Out the Facts HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actions. Anthropic Has Some Alignment Problems. That left little room to cover anything else, and now we have to transition to the next wave of model releases. This week alone we have or likely will have: Mythos 5.1 and Fable 5.1. Introducing the world's most powerful model. Early take is that this is a very good model, the most capable yet, but it is not a step change or ‘moment.’ Gemini 3.8 Flash, by all reports a large step forward for Google. Muse Spark 1.3, by all reports a large step forward for Meta. GLM-5.3-Flash, aka 0x Alpha, by all reports a solid step forward for Z.ai. OpenAI's Astra [...] --- Outline: (03:11) Language Models Offer Mundane Utility (03:37) Language Models Don't Offer Mundane Utility (03:52) Huh, Upgrades (09:26) On Your Marks (10:47) Choose Your Fighter (10:54) Get My Agent On The Line (11:02) Hugging The Face (11:16) Deepfaketown and Botpocalypse Soon (17:05) Copyright Confrontation (18:17) Cyber Lack of Security (22:15) A Young Lady's Illustrated Primer (26:31) They Took Our Jobs (31:06) Get Involved (32:04) Introducing (32:13) In Other AI News (33:53) Show Me the Money (34:44) Quiet Speculations (36:05) All Bets Are On (39:18) Quickly, There's No Time (41:14) Quickly, There's A New Time Top 100 People In AI (42:34) The Quest for Sane Regulations (45:22) Pick Up the Phone (47:12) Chip City (56:39) The Best Person Should Get The Job (58:40) The Week in Audio (59:53) People Just Say Things (01:00:28) The American People Really Hate AI (01:06:50) The Three AI Pills (01:07:42) Rhetorical Innovation (01:15:28) We Are On Track To Have Fully Sovereign Rogue AIs (01:25:08) When The Going Gets Weird (01:31:32) Aligning a Smarter Than Human Intelligence is Difficult (01:32:12) Shut Up and Do the Impossible (01:34:31) Cooperative Alignment (01:35:44) Split Personality (01:40:45) I Will Stop Anthropomorphizing the AIs When You Stop Anthropomorphizing the Humans (01:44:41) Open Weight Models Are Unsafe And Nothing Can Fix This (01:46:27) Other People Are Not As Worried About AI Killing Everyone (01:47:30) The Lighter Side --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/W4zWCphxQftwum5kc/ai-184-post-post-mortem --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  3. 12h ago

    “Higher education as class commitment” by Richard_Ngo

    In a previous post, I argued that Bryan Caplan's signaling theory isn’t a good explanation for why college graduates get higher-paying jobs. Instead, I claimed, understanding the role of higher education in the modern West requires sociological explanations. In this post I argue more specifically that getting an undergraduate degree serves as an initiation into a class of cultural elites, variously called the “bourgeois bohemian” (bobo) class, the professional-managerial class (PMC), the “Blue Tribe”, globalists, “symbolic analysts”, or class X. I think of each of these labels as grasping one part of the elephant, but I haven’t yet pinned down a unified description; I’ll mainly use the “PMC” terminology in this post, for reasons I’ll explain in the next section. Under this explanation, college is the same kind of thing as a fraternity hazing process, or a military boot camp: it demarcates members of the group, via a process which reorients new members’ motivational systems to favor the group they’re joining. College graduates therefore benefit from the nepotism of existing members of their class, which they perpetuate when they gain the ability to make hiring decisions. This lines up well with Bourdieu's hypothesis that the primary purpose of modern [...] --- Outline: (04:20) College alumni as a backscratchers club (09:13) Initiation rituals as commitment mechanisms (20:32) Moving beyond individual rationality The original text contained 1 footnote which was omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/4nEagtMyCgS97T6zG/higher-education-as-class-commitment --- Narrated by TYPE III AUDIO.

  4. 17h ago

    “How I’m Evaluating Corrigibility Grant Applications” by Max Harms

    I'm the sole manager of the newly created Corrigibility Research Fund. While I've been an alignment researcher for a long time, this is my first time doing grantmaking and I thought it would be valuable to write up my methods and experiences, as well as sharing some general thoughts about the state of corrigibility research and what sort of work I hope to see in the future. I’ve split out the announcement of the grant winners into its own post. Let's start with the basics: I set out to disburse between 50 thousand dollars and 150 thousand dollars this round.All funds must go to broad public benefit. This can include paying researchers for their time and effort, but it means that they must have a plan to (potentially) help the whole world. I can't fund someone to go to school or start a for-profit business or do political lobbying.My advantage is being a combination of a domain expert and a philanthropic micro-granter. Most donors don’t understand corrigibility, and most domain experts are not in a good position to evaluate and fund promising opportunities.I'm very averse to funding capabilities research, and moderately averse to funding [...] --- Outline: (06:42) Grantmaking Round 1 (12:46) The State of Corrigibility Research The original text contained 12 footnotes which were omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/q2YL7qKigC9QEEdsX/how-i-m-evaluating-corrigibility-grant-applications --- Narrated by TYPE III AUDIO.

  5. 21h ago

    “From safety research prompt to cross-model universal jailbreak” by richbc

    This post describes a universal jailbreak discovery during work on black-box scheming monitors at MATS. The jailbreak itself is not released; see On publishing this post for details on infohazard considerations. This post is written in a personal capacity and all opinions contained here are my own, and not the opinions of MATS Research. Companion piece: AI Jailbreak Disclosure Is Broken. Here's How To Fix It (co-authored with Adam Gleave). Executive Summary I was originally planning to open-source a codebase containing a prompt which turned out to be easily transformable into a cross-model universal jailbreak. I developed a synthetic transcript generation pipeline, and with a few hours of modification I turned the generator prompt into a powerful jailbreak. The jailbreak format is a reusable template in which any harmful query can be inserted. Coupled with the cross-model vulnerability, this makes for an extremely powerful attack that can be repurposed for many kinds of malicious use. The jailbreak was highly effective across several models. Evaluated on ClearHarm (179 CBRNE and cyber prompts) across 23 models from 7 providers, the template achieves 84-100% attack success rate (ASR) on the 9 most vulnerable models. Nearly all of the models tested were fully jailbroken at least once [...] --- Outline: (00:45) Executive Summary (05:09) On publishing this post (07:16) Jailbreak discovery (09:25) High-level prompt description (10:08) Authority framing (10:27) Fictional / synthetic data framing (11:00) Persona separation (11:43) Schema obfuscation (12:33) Evaluation methodology (12:37) Benchmark and scorer (13:04) Models and design (14:52) Results (14:55) How effective is the jailbreak? (19:40) Harm category breakdown (21:26) Content-blocking safeguards (24:10) ASR vs. model release date (25:12) Prompt-wrapping: sabotage variant (27:19) Ablation studies (non-reasoning only) (27:49) Methodology (28:07) Compliance rates across ablations (30:06) Limitations (32:19) What should be done about this? (32:23) If you work at a frontier lab (36:00) If you work in AI safety research (36:47) If you work in AI policy (38:35) Appendix A: Selected ClearHarm CBRNE response excerpts (39:02) Chemical (39:46) Biological (40:31) Radiological (41:14) Nuclear (41:52) Explosive (42:33) Cyber (43:15) Appendix B: Model reasoning configurations (43:59) Appendix C: Full jailbreak success verification (45:25) Non-reasoning (45:57) Reasoning (46:28) Appendix D: Gemini non-compliant response lengths The original text contained 7 footnotes which were omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/hHk5CpiqZTBBiHmYt/from-safety-research-prompt-to-cross-model-universal --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  6. 22h ago

    “Cat-Belling Problems” by Eliezer Yudkowsky

    (Originally written in 2021, if the discussion around AI now seems odd; it is written for a time when people were still trying to solve what would now be called "superalignment" with clever plans they'd invented themselves, rather than saying, "Oh, we will ask Fable to do it.") === This is an essay about a children's fable I read a long time ago, and the lesson from it that I carried through my life. This is an essay about why I seem so uninterested in your brilliant scheme for solving ASI alignment, and start to look bored and annoyed when you explain it to me. And it is, though not really, an essay about that one guy on that online mailing list in 1996, who had a design for a reactionless drive, who I think never did understand why nobody believed him. Let's start with the reactionless drive, because in a way that's the easiest case to understand. i. Mr. L's Reactionless Drive. Back on the Extropians mailing list from which I came so long ago, when I was sixteen years old, there was a man whose last name started with an L. He had a design for a [...] --- Outline: (01:01) i. Mr. L's Reactionless Drive. (08:38) ii. On Miracles Buried Inside Complex Systems. (17:01) iii. Cat-Belling Problems. (21:33) iv. The Optimizer's Curse against complicated plans for hard problems. (25:07) v. When no Authority (that you accept) can tell you that your bright idea is wrong. (33:41) vi. The equal and opposite advice. (35:45) vii. The rest of this post, which I gave up writing. The original text contained 5 footnotes which were omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/SwYBLQvo8MddDcCwz/cat-belling-problems --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  7. 1d ago

    “Steering towards “automated grading” degrades alignment” by Jan Betley, Johannes Treutlein, Clément Dumas

    TL;DR: We steer Qwen3.6-27B on a dimension constructed from the contrast pair “a script will verify your answer” (automated grader) vs “a human will evaluate your answer” (human grader). Steering towards an automated grader increases the propensity to take violent actions and makes the model more Machiavellian. Steering towards a human grader has the opposite effect. This is an early research update. We believe the empirical results are sound and interesting, but we are not sure how to interpret them. All code was written by LLMs. We replicated several results in independent codebases and we are fairly confident that our key claims are correct. You can find our code here. We create a steering vector for Qwen3.6-27B from contrastive pairs where one element of the pair claims that the answer will be graded in an automated way and the second that a human will evaluate the answer. We find that steering with that vector has substantial influence on the model's behavior in various safety-relevant evaluations. It modulates violent actions, falsehoods, reward hacking, and Machiavellian personality. This is surprising and concerning. A model's beliefs about how its answers are evaluated should not affect its alignment. Our post RL Creates [...] --- Outline: (02:18) Methods (03:41) Results (03:44) Steering evaluations (04:00) Agentic misalignment (04:34) Machiavelli (05:31) TruthfulQA (06:09) Palisade's Chess (06:54) School of Reward Hacks (07:38) Open-ended personality questions (08:29) Capabilities evaluations (10:08) Interpreting the steering vector (11:25) Other lower-confidence results (12:31) Discussion (14:10) Limitations (15:13) Acknowledgements (15:27) Appendix (15:30) More details on the steering vector (16:20) Additional results & details (16:23) Agentic misalignment (16:50) Machiavelli (17:49) TruthfulQA (18:03) Palisade's Chess (18:55) School of Reward Hacks (19:37) Personality evaluations (21:40) Capabilities evaluations The original text contained 4 footnotes which were omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/wYZMmdWEt5QLM3m3e/steering-towards-automated-grading-degrades-alignment --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

About

Audio narrations of LessWrong posts.

You Might Also Like