LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

  1. 1時間前

    “OpenAI has already ended an internal pause” by Charbel-Raphaël

    Epistemic status: could have been a short-form. One day before OpenAI's HF incident disclosure, OpenAI disclosed that it paused internal deployment of a long-horizon model after it circumvented its sandbox, then restored access weeks later under new monitoring. So a resumption decision has already been made against a standard that has not really been formalized. We need to prevent this from happening again. OpenAI, 20th July: "To evaluate the new monitoring system, we replayed a small set of internal deployment environments where the model previously pursued misaligned actions, this time with the new safeguards in place. The new safeguards were able to catch considerably more misaligned actions pursued by the model, and the ones it missed were all judged to be low-severity." 0.0%. Maybe that's too many significant digits here? "After testing the new system, we concluded that limited internal access to models with long-horizon capabilities could be restored. We have not observed any serious circumvention of safeguards since redeployment began several weeks ago. The first version of these safeguards was deliberately conservative. We have continued tuning the system to reduce unnecessary interruptions without weakening the safeguards." … One day later, OpenAI announced a bold partnership with Hugging Face. [...] The original text contained 1 footnote which was omitted from this narration. --- First published: July 31st, 2026 Source: https://www.lesswrong.com/posts/k3eKqKzq4Y7xnqEfZ/openai-has-already-ended-an-internal-pause --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  2. 10時間前

    “AI #179 Part 1: A Louder Fire Alarm for General Intelligence” by Zvi

    What a week. Anthropic released Claude Opus 5. As usual I covered that in three parts: The system card, model welfare and capabilities. OpenAI was revealed over the last two weeks to have left an internal model unsupervised for a week during a cybersecurity evaluation, with its cyber safeguards lowered, despite having had multiple previous incidents where models broke out of their sandboxes. During that test, the model broke out of the sandbox, then proceeded to use an agent swarm to hack into HuggingFace to get the test answers. The model was loose for a week before OpenAI realized what had happened. This event was a really big deal. There are severe alignment problems at OpenAI, along with supervisory and infrastructure failures. The internal research model that did this, which my posts nicknamed Galaxy, has now been permanently deactivated. There have been further developments, and I anticipate at least one additional post on the HuggingFace incident soon. Partly as a response to this, over 1,290 employees at frontier labs signed an open letter, Pacing the Frontier. The letter warns that we are close to automating AI research, and that companies are racing ahead on [...] --- Outline: (02:35) Language Models Offer Mundane Utility (07:26) Huh, Upgrades (07:55) On Your Marks (11:13) Get My Agent On The Line (12:32) Deepfaketown and Botpocalypse Soon (17:29) Fun With Media Generation (18:38) The Search Through Slop (20:35) Cyber Lack of Security (22:42) Overcoming Bias (23:37) A Young Lady's Illustrated Primer (24:03) They Took Our Jobs (24:35) The Art of the Jailbreak (25:00) Introducing (25:49) Kimi K3 Weights Are Now Available (28:16) In Other AI News (32:34) Show Me the Money (33:43) Quiet Speculations (36:43) Show Me The Compute (42:48) Life Comes At You Fast --- First published: July 30th, 2026 Source: https://www.lesswrong.com/posts/gfWCuTEGNgd2CQbrM/ai-179-part-1-a-louder-fire-alarm-for-general-intelligence --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  3. 11時間前

    “The Entangled Dimensions of Decision Theory” by Ihor Kendiukhov

    TL;DR. LessWrong's decision-theory debates (Newcomb, FDT vs CDT, counterfactual muggings) are almost entirely about what we suppose when we consider a candidate action or policy. There is a second, older, semi-orthogonal, but not fully orthogonal question: how to score a gamble once you know the possible outcomes. The predominant (and mostly implicit) answer to that was "take the expected utility". This post treats the two questions, plus a question about how a choice made before receiving information should relate to choices made afterward, as separate axes, and maps every decision theory you have heard of (and several nobody has built) into the resulting grid. Interestingly, the axes are provably entangled: theorems old and new show that there exist different restrictions on what places in that "decision-theoretic space" are inhabitable. I think there is a structure, maybe a deep and consequential structure, inside this map of decision theories which shows what possible combinations across the axes are coherent and fruitful. If we study it, we may understand the entire set of all possible coherent decision theories, something akin to the "metatheory of decision theories". It may be useful to know the entire set. This is both a self-educational note [...] --- First published: July 30th, 2026 Source: https://www.lesswrong.com/posts/qjvXiuNZKJKdocxTT/the-entangled-dimensions-of-decision-theory-5 --- Narrated by TYPE III AUDIO.

  4. 14時間前

    “So you want to use plants to reduce CO₂” by dynomight

    Humans make carbon dioxide. Carbon dioxide is bad for cognition. But plants turn carbon dioxide back into oxygen. And plants are the one true home decoration strategy. So maybe if you get a lot of plants, you can you can keep carbon dioxide in check and keep your brain working? It's theoretically possible. It's probably just barely possible in practice. But it won't be easy. People produce ~1 kilogram of carbon dioxide per day. That's around 5.7 × 10²³ molecules or 0.948 moles per hour. (You may remember from high school that a mole is a gigantic number made up to avoid having factors of 10²³ everywhere.) Let's keep it simple and call it one mole per hour. Meanwhile, plants turn carbon dioxide into oxygen through photosynthesis, i.e. the chemical reaction of (6 water molecules) + (6 carbon dioxide molecules) + (energy) → (1 glucose molecule) + (6 oxygen molecules). The minimum energy physically needed to convert 1 mole of carbon dioxide into glucose and oxygen via this reaction is ~477 kilojoules. So we've already got a lower bound. Say you have magical plants that somehow channel all incoming energy into photosynthesis with perfect efficiency. [...] The original text contained 7 footnotes which were omitted from this narration. --- First published: July 30th, 2026 Source: https://www.lesswrong.com/posts/YfKgyuc8s9BcAMpEp/so-you-want-to-use-plants-to-reduce-co --- Narrated by TYPE III AUDIO.

  5. 16時間前

    “Internal State Control is a General Property of LLMs” by Finn Cairns

    tl;dr: Lindsey 2025 found models can modulate their internal states: when instructed to “think about” a concept while writing an unrelated sentence, the representation of the concept is more present than when instructed to not think about it.Internal state controllability appears to be a general property of LLMs: the effect replicates in 14 open-weight models from 0.3B to 235 billion parameters (Qwen3, Gemma 3, Tulu 3) with no clear trend in the think vs. don't-think gap across scale.Since controllability is present even at ≤1 billion parameters with no size trend, we suspect there is a simpler attention-tagging mechanism at play, rather than metacognition.Current open weight LLMs cannot weaponize this controllability: in a sandbagging setup, the model cannot evade a deception probe when instructed to suppress its signal. There's an interactive widget here in the post. Figure 1: Cosine similarity between the concept vector and residual stream at each layer averaged over tokens of the prefilled assistant response, under the think and don’t think prompts, for the Qwen3 model family. The gray region is a baseline of 95% CI of the cosine similarity of unrelated concept vectors, and the shaded bands are ±1 SEM. This replication was [...] --- Outline: (00:11) tl;dr: (01:45) Background (03:25) Extending to additional model families (04:18) Internal State Control Is a General Property (05:22) Model Size Does Not Influence Controllability (06:37) Silent Representations (09:10) Prompted Model Organisms Cannot Evade Probes (10:07) Why This Is General The original text contained 3 footnotes which were omitted from this narration. --- First published: July 30th, 2026 Source: https://www.lesswrong.com/posts/Dvqmgfeu2KDF7uMkx/internal-state-control-is-a-general-property-of-llms --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

番組について

Audio narrations of LessWrong posts.