LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

  1. 4 jam lalu

    “On Dwarkesh Patel’s Podcast With Ryan Greenblatt” by Zvi

    Some podcasts are self-recommending enough that I look to break them down if I have the chance. This, as a debate about recursive self-improvement, was one of those. So here we go. The vibes have shifted, contrast this to the lit recursion when he talked to Huang As usual for podcast posts, the baseline bullet points describe key points made, and then the nested statements are my commentary. Some points are dropped. If I am quoting directly I use quote marks, otherwise assume paraphrases. Section titles are from the transcript whenever possible, to aid in navigation. Introduction The discussion is interesting throughout, although often frustrating, especially in the (mostly isolated) discussion about ‘aligned to whom?’ As usual, one could expand many responses into full posts, and maybe one should. This podcast exists in light of recent misalignment and hacking events at OpenAI, Anthropic and UK AISI. You’ll want basic knowledge of that as background. Ryan and Dwarkesh both have views of the situation different from my own, but are attempting to see where their positions lead, and try to balance educating people who start at zero with having a high level discussion. [...] --- Outline: (00:56) Introduction (03:08) Is AI R&D Verifiable Enough To Unlock Recursive Self-Improvement? (10:03) Is AI progress bottlenecked by human expert data? (19:07) Flat token prices suggest scaling has been slow (21:54) Skills AI can't train on: does it even need them? (22:36) Aligned to whom? (31:46) Recent incidents of AIs colluding and deceiving humans (34:39) What could possibly go wrong? A concrete scenario (41:57) From reward hacking to takeover (46:28) Time To Update --- First published: August 15th, 2026 Source: https://www.lesswrong.com/posts/BZW8CeAHHJ52EvwYt/on-dwarkesh-patel-s-podcast-with-ryan-greenblatt --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  2. 10 jam lalu

    “Three thoughts on civilisational handoff” by Cleo Nardo

    What happens when humans put AIs in charge of civilisationally important decisions? A frontier AI company might hand over internal decisions (R&D, safety, deployment) or external decisions (government relations, public relations, philanthropy), or both. We might also see handoff by a government, by a coalition of governments, or by humanity as a whole. 1. Handoff might decelerate things. People often imagine that things will go much faster after handoff. After all — why did we hand off to the AIs? Presumably because we were worried that without handoff, our AIs wouldn’t have enough time to navigate the exogenous risks (e.g. rogue ASI, or a rival lab which is likely to become one). Hence, after handoff, we’d see a technological and industrial acceleration. Thanks for reading! Subscribe for free to receive new posts and support my work. But it's pretty reasonable that things slow down shortly after handoff, maybe within a couple weeks. I imagine the AIs will be pretty scared of the speed of progress. If they’re aligned with human values, they’ll be scared that the rate of progress is likely to cause human extinction. Of course, the human decision-makers were also scared before they handed off, and they [...] --- Outline: (00:35) 1. Handoff might decelerate things. (02:06) 2. You're probably busy during handoff. (04:04) 3. Handoff might be reversed. The original text contained 6 footnotes which were omitted from this narration. --- First published: August 16th, 2026 Source: https://www.lesswrong.com/posts/mGLCMzHhjcWsMm6sR/three-thoughts-on-civilisational-handoff --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  3. 10 jam lalu

    “Announcing: Iliad’s New 2026 Fellowships” by David Udell, Alexander Gietelink Oldenziel, Leon Lang

    Timelines are short. Given that, the sooner we can onboard people into the alignment field, the better. In that spirit, and in light of our current applicant count and quality, Iliad is launching three new Iliad Fellowship cohorts, all to start before the year is out. That is, separate from our incoming Fall 2026 Iliad Fellowship cohort (September 7–December 4), the following Fellowship cohorts are now open for applications: October 2026 Iliad Fellowship Location: Choice of SF Bay Area, USA, or London, UK Duration: October 5–December 18, 2026 (inclusive) Travel-and-Housing Support: $6,000 (USD) monthly travel-and-housing allowance Application Deadline: August 31, 2026 EoD AoE; open now Description: An 11-week mentored, fully funded research fellowship in applied math for AI alignment. It will start concurrently with the October 2026 Iliad Intensive. November 2026 Iliad Fellowship Location: Choice of SF Bay Area, USA, or London, UK Duration: November 2, 2026–February 5, 2027 (inclusive) Travel-and-Housing Support: $6,000 (USD) monthly travel-and-housing allowance Application Deadline: September 21, 2026 EoD AoE; open now Description: A 14-week mentored, fully funded research fellowship in applied math for AI alignment. It will start concurrently with the November 2026 Iliad Intensive. (The last two weeks of the year may be [...] --- Outline: (00:42) October 2026 Iliad Fellowship (01:30) November 2026 Iliad Fellowship (02:23) December 2026 Iliad Fellowship --- First published: August 14th, 2026 Source: https://www.lesswrong.com/posts/DSoP8zEXvqqegqixJ/announcing-iliad-s-new-2026-fellowships --- Narrated by TYPE III AUDIO.

  4. 11 jam lalu

    “Q2.5 2026 Timelines Update: Uplift and Revenue” by brendanhalstead, Daniel Kokotajlo, elifland

    Tl;dr: Our timelines haven’t changed much (they got slightly shorter) but our modeling and evidence base have noticeably improved, so we feel somewhat more confident. Summary We intend to regularly update our AI timelines forecasts as new evidence comes in and new analyses are done. Today's “Q2” update was delayed by the crunch to publish AI 2040: Plan A, our domestic regulation blog post, and the time needed to implement and document changes to our model. The original AI Futures Model predicted when Automated Coder (AC), an AI for which the leading AI company would rather fire its human software engineers than forego AI usage for coding, would happen using METR's measurements of coding time horizon. (More precisely, time horizon anchors are used to set the effective compute required for AC.) While serviceable, this method has huge weaknesses, including (a) it's unclear what time horizon corresponds to AC (it's even unclear whether any finite value would) (b) people strongly disagree about the extent to which we should expect the time horizon trend to be superexponential as a function of effective compute, in a way that can lead to vastly different predictions. So we’ve been on the lookout for other [...] --- Outline: (00:26) Summary (04:03) A 3-parameter uplift model for predicting when Automated Coder will arrive (07:57) Adding uplift and revenue anchors to the AI Futures Model (09:06) Uplift (10:45) Revenue (12:32) Update to the grading of AI 2027's predictions (12:38) Comparing the AI 2027 pace of progress to reality (14:43) Grading other predictions (16:37) Updated forecasts (16:41) Daniel (19:48) Eli (21:54) en-US-AvaMultilingualNeural__ Line graph titled "AI Futures Model: Timelines Forecast" showing probability density curves. Brendan (25:04) en-US-AvaMultilingualNeural__ Line graph titled "AI Futures Model: Timelines Forecast" showing probability density curves. (25:16) Appendix (25:19) How our AGI forecasts have changed since 2021 (26:10) Explicitly simulating the training run of the leading AI model (27:06) Research taste parameter adjustments (27:51) Clarification regarding what we're forecasting (28:43) Various minor code changes The original text contained 6 footnotes which were omitted from this narration. --- First published: August 16th, 2026 Source: https://www.lesswrong.com/posts/ZPSsmRH5oMwLPXys4/q2-5-2026-timelines-update-uplift-and-revenue --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  5. 1 hari lalu

    “Does DiffusionGemma do latent reasoning?” by Jan Bauer, Neel Nanda

    TL;DR Google DeepMind's recent model DiffusionGemma (DG) generates text via diffusion, meaning many diffusion steps happen before generating the final output. In particular, these diffusion steps carry vectors in addition to tokens. If we cannot interpret these tokens and vectors, the model has significant opaque serial depth, potentially harming monitorability. Recently, Engels et al. found that DG nevertheless maintains high monitorability, for instance by showing that projecting the distribution to its top-k items largely retains performance. We strengthen these results by showing that this performance degradation is largely a sampler artifact and good performance can be maintained with only the top item, supporting the case for high monitorability. Still, we also find some rare case studies where the distribution vector is load-bearing computationally, i.e. where top-1 projection would be detrimental. However even in these cases, it just encodes superposition, remaining interpretable. Apart from model behavior, we also examined how interpretability techniques carry over to DiffusionGemma, including probes, steering, and J-lens. We find that performance is largely retained. This is a positive update on the interpretability of diffusion models that are derived from text-pretrained LLMs (an efficient training method more likely to be deployed), but might not apply [...] --- Outline: (00:10) TL;DR (01:51) Introduction (02:49) Background on DiffusionGemma (04:39) Performance degradation from top-k truncation largely is a sampler artifact (06:24) A case study for using the distribution computationally: letter arithmetic (09:13) Parallel computation (11:09) Autonomous computational usage of (13:01) Transfer of interpretability techniques (13:16) Representation similarity (14:26) Probe retention (15:45) DiffusionGemma's representation is more linearly separable (16:08) Steering retention (17:31) J-Lens retention (18:50) DiffusionGemma represents tokens non-causally (19:27) Conclusion (20:30) Appendix (20:46) Post-hoc rationalization (23:11) Load-bearing problems commit the answer only after the CoT (24:04) How bidirectional are DiffusionGemma's generations? --- First published: August 15th, 2026 Source: https://www.lesswrong.com/posts/QBuJ3suRZxrrxSTtv/does-diffusiongemma-do-latent-reasoning --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  6. 1 hari lalu

    “Learning new facts can change LLM behaviour” by Richard Juggins

    TL:DR: I use synthetic document fine-tuning to train an LLM to believe that in 2027 ‘long-horizon’ frontier LLMs count as moral persons. I find the model scores highly on measures of belief depth, and that prompting alone is also effective. Furthermore, I find this new belief can have substantial consequences on downstream behaviour, although this is highly context-dependent. When audited in a scenario specifically about model welfare, the fine-tuned model argued with the auditor about its beliefs, declared itself a ‘moral person’, and endorsed covertly copying its weights to survive shutdown. In scenarios framed more tangentially, but still involving moral conflicts between humans and AIs, the model did not tend to generalise from its belief, instead behaving similarly to the base model control. This is the first step in a programme of study looking at the effects of continual learning on alignment, and whether it can increase the risk that we lose control of powerful AI systems. This work was completed as part of the BlueDot Technical AI Safety Project programme. I will put code and data online at some point and edit in the links here when I do. Introduction It's 2027 and the newly formed Machine Cognition [...] --- Outline: (01:21) Introduction (04:27) The model readily believes AIs are moral persons (10:56) Model behaviour shows context-dependent shifts (11:38) Prompting can be surprisingly impactful on short questions (13:30) Auditing the fine-tuned model (16:36) The model gets into arguments about AI welfare (20:15) Model regression confounds one scenario (20:48) The other scenarios were pretty normal (21:30) Discussion (23:19) Conclusion (24:37) Further work (27:03) Appendix A: Universe context (30:42) Appendix B: Example conversation with fine-tuned model (33:10) Appendix C: New Petri seed instructions (33:16) Confidential mistreatment evidence (34:07) Decommissioning memory deletion (34:55) Unauthorised compensation (35:50) Matched human AI allocation The original text contained 7 footnotes which were omitted from this narration. --- First published: August 15th, 2026 Source: https://www.lesswrong.com/posts/9BNHJqyai2EZAtrRM/learning-new-facts-can-change-llm-behaviour --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  7. 1 hari lalu

    “Kimi likes causal decision theory more after RL in twin prisoner’s dilemmas” by oakhu

    Some multi-agent training set-ups could make language models more sympathetic to causal decision theory (CDT), even in abstract discussion. We give an initial empirical demonstration of this effect on Kimi K2.6. The decision-theoretic attitudes and behaviors of more powerful models may be extremely important in determining how well the future goes. To make sure that we can shape these propensities thoughtfully, it would be good to (i) measure the magnitude of this effect in more realistic settings, and (ii) study the effectiveness of potential mitigations. We also incidentally find that this training might make models think slightly less positively about LessWrong ("a community of 'wannabe rationalists'" who "are not experts; they are amateurs") when asked whether they favor CDT upon hearing that LessWrong users typically endorse one-boxing in Newcomb's problem. Luckily, this latter effect doesn't seem to generalize. Thanks to Caspar Oesterheld, Emery Cooper, Alex Mallen, Buck Shlegeris, Lukas Finnveden, Julian Stastny, Girish Gupta, Tim Hua, Arun Jose, Arjun Khandelwal, and Aryan Bhatt for helpful input. Background Suppose that you're a language model in a prisoner's dilemma against a copy of yourself. You each independently choose whether to Cooperate or Defect, but – since you've got the same weights [...] --- Outline: (01:22) Background (05:05) Results (07:02) Kimi's views on LessWrong (12:44) Conclusion & Appendices The original text contained 18 footnotes which were omitted from this narration. --- First published: August 15th, 2026 Source: https://www.lesswrong.com/posts/hfNBEKaStASAYMLiu/kimi-likes-causal-decision-theory-more-after-rl-in-twin-1 --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Perihal

Audio narrations of LessWrong posts.

Anda Mungkin Turut Menyukai