LessWrong (30+ Karma)

LessWrong

0.0 (0)
科技
一日一更

Audio narrations of LessWrong posts.

2小时前

“Training fails to elicit subtle reasoning in current language models” by mishajw, Fabien Roger, Hoagy, gasteigerjo, Joe Benton, Vlad Mikulik

While recent AI systems achieve strong performance through human-readable reasoning that should be simple to monitor (OpenAI, 2024, Anthropic, 2025), we investigate whether models can learn to reason about malicious side tasks while making that reasoning appear benign. We find that Sonnet 3.7 can learn to evade either a reasoning monitor, by persuading the monitor that a blatant backdoor is benign, or an output-only monitor, by devising sophisticated backdoors that the output-only monitor doesn’t detect. But when trained to evade both reasoning and output-only monitors, Sonnet 3.7 is unable to use reasoning to improve its backdoor success rate without triggering a reasoning monitor. Like previous work (Baker et al., 2025, Emmons et al., 2025), our results suggest that reasoning monitors can provide strong assurance that language models are not pursuing reasoning-heavy malign side tasks, but that additional mitigations may be required for robustness to monitor persuasion. Figure 1: We trained [...] --- First published: October 9th, 2025 Source: https://www.lesswrong.com/posts/MmuyzfsaNrSvRCsFk/training-fails-to-elicit-subtle-reasoning-in-current --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

8 分钟
7小时前

“‘Yes, and—’ Requires the Possibility of ‘No, Because—’” by Zack_M_Davis

Scott Garrabrant gives a number of examples to illustrate that "Yes Requires the Possibility of No". We can understand the principle in terms of information theory. Consider the answer to a yes-or-no question as a binary random variable. The "amount of information" associated with a random variable is quantified by the entropy, the expected value of the negative logarithm of the probability of the outcome. If we know in advance of asking that the answer to the question will always be Yes, then the entropy is −P(Yes)·log(P(Yes)) − P(No)·log(P(No)) = −1·log(1) − 0·log(0) = 0.[1] If you already knew what the answer would be, then the answer contains no information; you didn't learn anything new by asking. In the art of improvisational theater ("improv" for short), actors perform scenes that they make up as they go along. Without a script, each actor's choices of what to say and [...] The original text contained 2 footnotes which were omitted from this narration. --- First published: October 9th, 2025 Source: https://www.lesswrong.com/posts/Pwg7nmjkx8mxmE6gF/yes-and-requires-the-possibility-of-no-because --- Narrated by TYPE III AUDIO.

7 分钟
11小时前

“Stars are a rounding error” by Algon

Notes on some interesting factoids I learnt from Anders Sandberg's draft book, Grand Futures. "Starlight is heavier than worlds" - Anders Sandberg Looking at the energy density of stuff in the universe, we find a few surprising, and not so surprising, facts. First, the obvious: baryonic matter itself is a rounding error, contributing 4.5% of the energy of the universe. Nine tenths of those sweet, sweet baryonic numéraire are stuck in the warm plasma floating between galaxies. About half the remainder forms the stuff of stars. Planets don't even match a thousand of the contribution of stars to the energy density of the universe. Somewhat surprisingly, supermassive black holes have a contribution. Regardless, the fact remains that planets are a rounding error to a rounding error to a rounding error of the energy of dark matter and energy. Even starlight contains more energy. So in a [...] --- First published: October 9th, 2025 Source: https://www.lesswrong.com/posts/mzifm6wePKfnFTAaB/stars-are-a-rounding-error --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

6 分钟
14小时前

“Training Qwen-1.5B with a CoT legibility penalty” by Fabien Roger

Audio note: this article contains 102 uses of latex notation, so the narration may be difficult to follow. There's a link to the original text in the episode description. I tried training Qwen2.5-1.5B with RL on math to both get correct answers and have a CoT that doesn’t look like human-understandable math reasoning. RL sometimes succeeds at hacking my monitor, and when I strengthen my monitor, it fails at finding CoT that are both illegible and helpful, even after training for roughly 4000 steps (~1B generated tokens). Exploring into obfuscated reasoning is hard! These results were also released in the Appendix of Training fails to elicit subtle reasoning in current language models. Motivation Chain-of-Thoughts (CoTs) can help reason for many more serial steps than there are layers in a Transformer. But one worry is that LLMs might hide their real reasoning in a plausible benign CoT. Previous [...] --- Outline: (00:58) Motivation (01:45) Methods (03:33) Results (07:02) Limitations (07:52) Appendix (07:55) Hackable illegibility rater prompt (09:03) Stronger illegibility rater prompt (09:25) System prompt used for the policy (09:40) System prompt used for the no-CoT policy (09:48) Example generations --- First published: October 9th, 2025 Source: https://www.lesswrong.com/posts/WSKNmRxPnYdQnoNvt/training-qwen-1-5b-with-a-cot-legibility-penalty --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

10 分钟
16小时前

“At odds with the unavoidable meta-message” by Ruby

It is a truism known to online moderators[1] that when two commenters are going back and forth in heated exchange, and one lays out rejoinders in paragraph after paragraph of dense text, then two things will have happened: Our careful communicator may or may not have succeeded at conveying well-reasoned insights hitherto unknown by his interlocutor that will change her mind. He will have communicated my hatred for you is at least this long. In all seriousness, words require effort and effort requires motivation. A lengthy message communicates whatever its contents are, but it inevitably communicates I was invested enough to write all this, typically with a connotation of strong emotion powering the keystrokes. The emotion isn't necessarily hatred. It could be a roiling anger, a neurotic anxiety, or a fervorous infatuation. My guess is that even if you weren't carrying around an explicit belief in [...] The original text contained 5 footnotes which were omitted from this narration. --- First published: October 10th, 2025 Source: https://www.lesswrong.com/posts/AgkuN8wLNevBkHusf/at-odds-with-the-unavoidable-meta-message --- Narrated by TYPE III AUDIO.

7 分钟
19小时前

“Towards a Typology of Strange LLM Chains-of-Thought” by 1a3orn

Intro LLMs being trained with RLVR (Reinforcement Learning from Verifiable Rewards) start off with a 'chain-of-thought' (CoT) in whatever language the LLM was originally trained on. But after a long period of training, the CoT sometimes starts to look very weird; to resemble no human language; or even to grow completely unintelligible. Why might this happen? I've seen a lot of speculation about why. But a lot of this speculation narrows too quickly, to just one or two hypotheses. My intent is also to speculate, but more broadly. Specifically, I want to outline six nonexclusive possible causes for the weird tokens: new better language, spandrels, context refresh, deliberate obfuscation, natural drift, and conflicting shards. And I also wish to extremely roughly outline ideas for experiments and evidence that could help us distinguish these causes. I'm sure I'm not enumerating the full space of [...] --- Outline: (00:11) Intro (01:34) 1. New Better Language (04:06) 2. Spandrels (06:42) 3. Context Refresh (10:48) 4. Deliberate Obfuscation (12:36) 5. Natural Drift (13:42) 6. Conflicting Shards (15:24) Conclusion --- First published: October 9th, 2025 Source: https://www.lesswrong.com/posts/qgvSMwRrdqoDMJJnD/towards-a-typology-of-strange-llm-chains-of-thought --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

18 分钟
20小时前

“I take antidepressants. You’re welcome” by Elizabeth

It's amazing how much smarter everyone else gets when I take antidepressants. It makes sense that the drugs work on other people, because there's nothing in me to fix. I am a perfect and wise arbiter of not only my own behavior but everyone else's, which is a heavy burden because some of ya’ll are terrible at life. You date the wrong people. You take several seconds longer than necessary to order at the bagel place. And you continue to have terrible opinions even after I explain the right one to you. But only when I’m depressed. When I’m not, everyone gets better at merging from two lanes to one. This effect is not limited by the laws of causality or time. Before I restarted Wellbutrin, my partner showed me this song. My immediate reaction was, “This is fine, but what if [...] --- Outline: (04:39) Caveats (05:27) Acknowledgements --- First published: October 9th, 2025 Source: https://www.lesswrong.com/posts/FnrhynrvDpqNNx9SC/i-take-antidepressants-you-re-welcome --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

6 分钟
21小时前

“Realistic Reward Hacking Induces Different and Deeper Misalignment” by Jozdien

TL;DR: I made a dataset of realistic harmless reward hacks and fine-tuned GPT-4.1 on it. The resulting models don't show emergent misalignment on the standard evals, but they do alignment fake (unlike models trained on toy reward hacks), seem more competently misaligned, are highly evaluation-aware, and the effects persist when mixing in normal data. Thanks to Aidan Ewart, Jack Kaunismaa, Abhay Sheshadri, Maxime Riché, Axel Ahlqvist, Niels Warncke, Daniel Tan, Carolyn Qian, and Kei Nishimura-Gasparian for helpful conversations, comments and/or feedback. This post is best viewed as an informal report on preliminary results done over a couple days, rather than a very polished analysis. Introduction Taylor et al finds that fine-tuning LLMs on harmless reward hacks causes generalization to unrelated misaligned behavior on the emergent misalignment (EM) evals. They constructed a fine-tuning dataset (School of Reward Hacks) of samples like this: There's a details box here with the title "Sample [...] --- Outline: (00:56) Introduction (03:17) Dataset (05:24) Emergent Misalignment Evals (07:34) Alignment Faking (16:29) Takeaways (18:28) How robust is this effect? The original text contained 11 footnotes which were omitted from this narration. --- First published: October 9th, 2025 Source: https://www.lesswrong.com/posts/HLJoJYi52mxgomujc/realistic-reward-hacking-induces-different-and-deeper-1 --- Narrated by TYPE III AUDIO. --- Images from the article: 5 out of 10." style="max-width: 100%;" />Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

22 分钟

查看全部 250 集

Audio narrations of LessWrong posts.

创作者

LessWrong
活跃年份

2023年 - 2025年
单集

250
分级

儿童适宜
节目网站

LessWrong (30+ Karma)

科技

科技

两周一更
科技

科技

一周一更
营养

营养

9月29日更新
科学

科学

10月3日更新
健康与健身

健康与健身

一周一更
科学

科学

6天前更新
健康与健身

健康与健身

一周一更

LessWrong (30+ Karma)

“Training fails to elicit subtle reasoning in current language models” by mishajw, Fabien Roger, Hoagy, gasteigerjo, Joe Benton, Vlad Mikulik

“‘Yes, and—’ Requires the Possibility of ‘No, Because—’” by Zack_M_Davis

“Stars are a rounding error” by Algon

“Training Qwen-1.5B with a CoT legibility penalty” by Fabien Roger

“At odds with the unavoidable meta-message” by Ruby

“Towards a Typology of Strange LLM Chains-of-Thought” by 1a3orn

“I take antidepressants. You’re welcome” by Elizabeth

“Realistic Reward Hacking Induces Different and Deeper Misalignment” by Jozdien

关于

信息

你可能还喜欢

LessWrong (30+ Karma)

单集

“Training fails to elicit subtle reasoning in current language models” by mishajw, Fabien Roger, Hoagy, gasteigerjo, Joe Benton, Vlad Mikulik

“‘Yes, and—’ Requires the Possibility of ‘No, Because—’” by Zack_M_Davis

“Stars are a rounding error” by Algon

“Training Qwen-1.5B with a CoT legibility penalty” by Fabien Roger

“At odds with the unavoidable meta-message” by Ruby

“Towards a Typology of Strange LLM Chains-of-Thought” by 1a3orn

“I take antidepressants. You’re welcome” by Elizabeth

“Realistic Reward Hacking Induces Different and Deeper Misalignment” by Jozdien

关于

信息

你可能还喜欢