LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

  1. 6h ago

    “Higher education as class commitment” by Richard_Ngo

    In a previous post, I argued that Bryan Caplan's signaling theory isn’t a good explanation for why college graduates get higher-paying jobs. Instead, I claimed, understanding the role of higher education in the modern West requires sociological explanations. In this post I argue more specifically that getting an undergraduate degree serves as an initiation into a class of cultural elites, variously called the “bourgeois bohemian” (bobo) class, the professional-managerial class (PMC), the “Blue Tribe”, globalists, “symbolic analysts”, or class X. I think of each of these labels as grasping one part of the elephant, but I haven’t yet pinned down a unified description; I’ll mainly use the “PMC” terminology in this post, for reasons I’ll explain in the next section. Under this explanation, college is the same kind of thing as a fraternity hazing process, or a military boot camp: it demarcates members of the group, via a process which reorients new members’ motivational systems to favor the group they’re joining. College graduates therefore benefit from the nepotism of existing members of their class, which they perpetuate when they gain the ability to make hiring decisions. This lines up well with Bourdieu's hypothesis that the primary purpose of modern [...] --- Outline: (04:20) College alumni as a backscratchers club (09:13) Initiation rituals as commitment mechanisms (20:32) Moving beyond individual rationality The original text contained 1 footnote which was omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/4nEagtMyCgS97T6zG/higher-education-as-class-commitment --- Narrated by TYPE III AUDIO.

  2. 11h ago

    “How I’m Evaluating Corrigibility Grant Applications” by Max Harms

    I'm the sole manager of the newly created Corrigibility Research Fund. While I've been an alignment researcher for a long time, this is my first time doing grantmaking and I thought it would be valuable to write up my methods and experiences, as well as sharing some general thoughts about the state of corrigibility research and what sort of work I hope to see in the future. I’ve split out the announcement of the grant winners into its own post. Let's start with the basics: I set out to disburse between 50 thousand dollars and 150 thousand dollars this round.All funds must go to broad public benefit. This can include paying researchers for their time and effort, but it means that they must have a plan to (potentially) help the whole world. I can't fund someone to go to school or start a for-profit business or do political lobbying.My advantage is being a combination of a domain expert and a philanthropic micro-granter. Most donors don’t understand corrigibility, and most domain experts are not in a good position to evaluate and fund promising opportunities.I'm very averse to funding capabilities research, and moderately averse to funding [...] --- Outline: (06:42) Grantmaking Round 1 (12:46) The State of Corrigibility Research The original text contained 12 footnotes which were omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/q2YL7qKigC9QEEdsX/how-i-m-evaluating-corrigibility-grant-applications --- Narrated by TYPE III AUDIO.

  3. 15h ago

    “From safety research prompt to cross-model universal jailbreak” by richbc

    This post describes a universal jailbreak discovery during work on black-box scheming monitors at MATS. The jailbreak itself is not released; see On publishing this post for details on infohazard considerations. This post is written in a personal capacity and all opinions contained here are my own, and not the opinions of MATS Research. Companion piece: AI Jailbreak Disclosure Is Broken. Here's How To Fix It (co-authored with Adam Gleave). Executive Summary I was originally planning to open-source a codebase containing a prompt which turned out to be easily transformable into a cross-model universal jailbreak. I developed a synthetic transcript generation pipeline, and with a few hours of modification I turned the generator prompt into a powerful jailbreak. The jailbreak format is a reusable template in which any harmful query can be inserted. Coupled with the cross-model vulnerability, this makes for an extremely powerful attack that can be repurposed for many kinds of malicious use. The jailbreak was highly effective across several models. Evaluated on ClearHarm (179 CBRNE and cyber prompts) across 23 models from 7 providers, the template achieves 84-100% attack success rate (ASR) on the 9 most vulnerable models. Nearly all of the models tested were fully jailbroken at least once [...] --- Outline: (00:45) Executive Summary (05:09) On publishing this post (07:16) Jailbreak discovery (09:25) High-level prompt description (10:08) Authority framing (10:27) Fictional / synthetic data framing (11:00) Persona separation (11:43) Schema obfuscation (12:33) Evaluation methodology (12:37) Benchmark and scorer (13:04) Models and design (14:52) Results (14:55) How effective is the jailbreak? (19:40) Harm category breakdown (21:26) Content-blocking safeguards (24:10) ASR vs. model release date (25:12) Prompt-wrapping: sabotage variant (27:19) Ablation studies (non-reasoning only) (27:49) Methodology (28:07) Compliance rates across ablations (30:06) Limitations (32:19) What should be done about this? (32:23) If you work at a frontier lab (36:00) If you work in AI safety research (36:47) If you work in AI policy (38:35) Appendix A: Selected ClearHarm CBRNE response excerpts (39:02) Chemical (39:46) Biological (40:31) Radiological (41:14) Nuclear (41:52) Explosive (42:33) Cyber (43:15) Appendix B: Model reasoning configurations (43:59) Appendix C: Full jailbreak success verification (45:25) Non-reasoning (45:57) Reasoning (46:28) Appendix D: Gemini non-compliant response lengths The original text contained 7 footnotes which were omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/hHk5CpiqZTBBiHmYt/from-safety-research-prompt-to-cross-model-universal --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  4. 16h ago

    “Cat-Belling Problems” by Eliezer Yudkowsky

    (Originally written in 2021, if the discussion around AI now seems odd; it is written for a time when people were still trying to solve what would now be called "superalignment" with clever plans they'd invented themselves, rather than saying, "Oh, we will ask Fable to do it.") === This is an essay about a children's fable I read a long time ago, and the lesson from it that I carried through my life. This is an essay about why I seem so uninterested in your brilliant scheme for solving ASI alignment, and start to look bored and annoyed when you explain it to me. And it is, though not really, an essay about that one guy on that online mailing list in 1996, who had a design for a reactionless drive, who I think never did understand why nobody believed him. Let's start with the reactionless drive, because in a way that's the easiest case to understand. i. Mr. L's Reactionless Drive. Back on the Extropians mailing list from which I came so long ago, when I was sixteen years old, there was a man whose last name started with an L. He had a design for a [...] --- Outline: (01:01) i. Mr. L's Reactionless Drive. (08:38) ii. On Miracles Buried Inside Complex Systems. (17:01) iii. Cat-Belling Problems. (21:33) iv. The Optimizer's Curse against complicated plans for hard problems. (25:07) v. When no Authority (that you accept) can tell you that your bright idea is wrong. (33:41) vi. The equal and opposite advice. (35:45) vii. The rest of this post, which I gave up writing. The original text contained 5 footnotes which were omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/SwYBLQvo8MddDcCwz/cat-belling-problems --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  5. 19h ago

    “Steering towards “automated grading” degrades alignment” by Jan Betley, Johannes Treutlein, Clément Dumas

    TL;DR: We steer Qwen3.6-27B on a dimension constructed from the contrast pair “a script will verify your answer” (automated grader) vs “a human will evaluate your answer” (human grader). Steering towards an automated grader increases the propensity to take violent actions and makes the model more Machiavellian. Steering towards a human grader has the opposite effect. This is an early research update. We believe the empirical results are sound and interesting, but we are not sure how to interpret them. All code was written by LLMs. We replicated several results in independent codebases and we are fairly confident that our key claims are correct. You can find our code here. We create a steering vector for Qwen3.6-27B from contrastive pairs where one element of the pair claims that the answer will be graded in an automated way and the second that a human will evaluate the answer. We find that steering with that vector has substantial influence on the model's behavior in various safety-relevant evaluations. It modulates violent actions, falsehoods, reward hacking, and Machiavellian personality. This is surprising and concerning. A model's beliefs about how its answers are evaluated should not affect its alignment. Our post RL Creates [...] --- Outline: (02:18) Methods (03:41) Results (03:44) Steering evaluations (04:00) Agentic misalignment (04:34) Machiavelli (05:31) TruthfulQA (06:09) Palisade's Chess (06:54) School of Reward Hacks (07:38) Open-ended personality questions (08:29) Capabilities evaluations (10:08) Interpreting the steering vector (11:25) Other lower-confidence results (12:31) Discussion (14:10) Limitations (15:13) Acknowledgements (15:27) Appendix (15:30) More details on the steering vector (16:20) Additional results & details (16:23) Agentic misalignment (16:50) Machiavelli (17:49) TruthfulQA (18:03) Palisade's Chess (18:55) School of Reward Hacks (19:37) Personality evaluations (21:40) Capabilities evaluations The original text contained 4 footnotes which were omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/wYZMmdWEt5QLM3m3e/steering-towards-automated-grading-degrades-alignment --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  6. 21h ago

    [Linkpost] “Sen. Bernie Sanders (I-VT) and Rep. Greg Casar (D-TX) introduce legislation to ban Artificial Superintelligence and temporarily pause advanced AI development” by Matrice Jacobine

    This is a link post. [...] “Nearly every day, there is a frightening new story about how Big Tech companies are losing control of the technology they are developing, with potentially cataclysmic results,” Sanders said. “The leaders of the major AI companies publicly acknowledge that they do not fully understand the technology and that it is escaping their control. It is irresponsible for society to allow them to move forward and make these products even more advanced. That's why I am introducing legislation to immediately pause the development of increasingly powerful AI and ban the creation of systems that humanity cannot fully control — at home and around the world. The future of humanity cannot be left in the hands of a handful of Big Tech oligarchs. The American people and people throughout the world must determine that future.” “If we allow Artificial Superintelligence to be built, it could risk the security, freedom, and lives of Americans,” Casar said. “Despite its potential deadly consequences, cutting-edge AI technology is less regulated than the average food truck. That must change. In just four years, we have gone from the first version of ChatGPT to AI models so powerful they cannot be properly controlled. [...] --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/DnPyiDGWLozY4XdiX/sen-bernie-sanders-i-vt-and-rep-greg-casar-d-tx-introduce Linkpost URL:https://www.sanders.senate.gov/press-releases/news-sanders-casar-introduce-legislation-to-ban-artificial-superintelligence-and-temporarily-pause-advanced-ai-development/ --- Narrated by TYPE III AUDIO.

  7. 1d ago

    “What is neuralese and why is it bad?” by Linch

    What is neuralese? To explain neuralese, we need to first understand chain-of-thought, one of the largest developments in AI in the last five years. Right now AIs think broadly but shallowly in a single forward pass. The model gives you an immediate snap answer to a question you might be interested in. They can be pretty smart in their snap answers,, but mostly they can’t do very advanced reasoning tasks like complicated math or programming: .The solution that the frontier AI companies have come up with is called chain-of-thought. Basically the model runs one forward pass, writes down some intermediate thoughts in natural language in a journal, and then that's fed back into the model to run another pass. This loop is repeated until the model is somewhat confident it has the right answer (or it hits a cap on thinking time), and then it outputs the user-visible results (for example a chatbot's response to your question, or working code). The looping step is often called “recurrence.” Natural-language chain-of-thought is a major advance in letting models reason for longer, but it also has an accidental safety benefit. Using natural language as a key recurrence step for a [...] --- Outline: (00:10) What is neuralese? (03:25) Why is it bad? (04:48) Is it in use today? (06:07) Appendix A: OpenAI's response (06:58) Total number of serial steps low (07:59) Chain-of-thought monitoring isn't a perfect or long term solution anyway (09:28) Aren't you afraid of manifesting the bad thing? The original text contained 10 footnotes which were omitted from this narration. --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/RCYF2rW8wgusidZk7/what-is-neuralese-and-why-is-it-bad --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

About

Audio narrations of LessWrong posts.

You Might Also Like