LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

  1. 2 小時前

    “SOTA alignment assessments don’t strongly update us against misalignment” by Alexa Pan

    Anthropic concluded in the April Mythos Preview alignment risk update that the model "does not possess any unknown propensities that would increase alignment risk." The report argues that if Mythos Preview were coherently misaligned, it likely would have been detected by the assessment (following Anthropic, I will call this “reliability of the assessment”). While I agree with the report on the above bottom-line conclusions (substantially on priors), I think there are gaps in its argument which weaken the current assessment and might invalidate future assessments. In particular, the report often uses weak evidence to justify reliability. The report gives fairly weak experimental evidence for Mythos Preview having insufficient capabilities to evade monitoring. The model is plausibly often eval-aware and underelicited in the relevant capability evaluations. So, it might silently sandbag if coherently misaligned, or unintentionally underperform if otherwise misaligned. This limitation is important: one could argue that lack of covert capabilities for sophisticated sabotage (a subset of the capabilities I discuss here) is the single most load bearing argument in alignment risk reports.Authors of the report could have made calibrated guesses about Mythos Preview's covert capabilities, especially for covert sabotage, based on other factors despite the relatively weak [...] --- Outline: (03:14) How reliability fits into the overall safety argument (05:20) Reliability claims by AI companies (05:56) Reliability claims by external evaluators (06:32) Alignment assessments are less reliable than developers claim (07:15) 1: Measuring capabilities to covertly undermine alignment assessments (10:05) Issues with evaluation awareness (13:28) Issues with underestimating covert capabilities (16:34) Issues with sandbagging rule-out (19:10) 2: Stress-testing alignment assessments with auditing games (20:16) An auditing failure with Mythos (22:08) AuditBench results (24:04) 3: Conditioning on misalignment should make us think that certain covert capabilities are better than expected (26:04) Bottom line on the strength of current alignment assessments (28:52) Conclusion (29:29) Appendix: (29:32) Why I focus on motive / alignment assessments in alignment risk reports (30:44) Auditability vs. Trustedness (33:11) More reliability claims by developers and third party evaluators (33:27) Mythos Alignment Risk Update (34:41) Opus 4.6 Sabotage Risk Report (35:24) GPT 5.5 System card (36:21) Muse Spark system card (37:08) Mythos Alignment Risk Update, safety arguments against sandbagging (37:15) From the Mythos Alignment Update, §5.3.4, p. 24: (38:09) UK AISI evaluations for Opus 4.7 (39:28) Past auditing games by Anthropic (41:45) Anti-auditing capability measurements (43:12) Conditioning on coherent misalignment updates us on certain covert capabilities The original text contained 46 footnotes which were omitted from this narration. --- First published: July 31st, 2026 Source: https://www.lesswrong.com/posts/oirrSj3itFLSyscW8/sota-alignment-assessments-don-t-strongly-update-us-against --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  2. 4 小時前

    “AI #179 Part 2: Hearing The Fire Alarm” by Zvi

    This is a continuation of Part 1 from yesterday. The back portion of the update, as usual, deals with policy, rhetoric, risk and alignment. I had to include an extended discussion of the other open letter, the one about open weight models, but most of you can skip those sections entirely, which is why they are in italics in the Table of Contents. Table of Contents The Frontier Act. This likely deserves a full RTFB but I haven’t had the time. The Quest for Sane Regulations. Sam Altman goes to Washington. Leading the Future Never Changes. They also do not plan to apologize. Chip City. Do not ban the Chinese robots, that will only make things worse. The Week in Audio. Altman twice, the AI 2027 team. People Just Say Yay Open Weights. An open letter. Open Weights Frontier Models Are Unsafe And Nothing Can Fix This. People Just Say Things. Push The Magic Button. Not you can. But if you could. Rhetorical Innovation. Distinctions between different arguments. Joshua Achiam's Final Message Upon Leaving OpenAI. Never stop. Dear Dario and Amanda. Claude [...] --- Outline: (00:35) The Frontier Act (03:31) The Quest for Sane Regulations (10:31) Leading the Future Never Changes (12:22) Chip City (17:30) The Week in Audio (19:20) People Just Say Yay Open Weights (32:32) Open Weights Frontier Models Are Unsafe And Nothing Can Fix This (37:19) People Just Say Things (47:20) Push The Magic Button (50:49) Rhetorical Innovation (56:31) Joshua Achiam's Final Message Upon Leaving OpenAI (59:51) Dear Dario and Amanda (01:09:58) Other People Are Not As Worried About AI Killing Everyone (01:12:30) How To Contact Me (01:14:37) The Lighter Side --- First published: July 31st, 2026 Source: https://www.lesswrong.com/posts/CXeoAhNrAeWpvoyiF/ai-179-part-2-hearing-the-fire-alarm --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  3. 5 小時前

    “My Assessment of Compute Verification in Plan A (+ open questions)” by jacob_drori

    Overview These are my non-expert notes on the compute verification section of AIFP's Plan A. I cover interconnect limits, memory wipes, network taps + replay, and ZKPs. For the most part, the sections can be read independently. I restrict my attention to inference-only verification: ensuring that compute is used for inference, not training. For each method suggested by AIFP, I ask: How much can it slow down training?How much overhead does it add to inference?What sensitive information does it require adversaries to share with each other? AIFP estimates that the fraction of the world's compute that is unmonitored might be kept as low as 0.1% (this is the optimistic, low end of their 80% confidence interval). So my target for inference-only verification is to slow down training by 1000 times – any more hits diminishing returns as unmonitored compute dominates – with much less than 1000x overhead on inference and little sharing of secrets. I won’t discuss how much a 1000x reduction in effective training compute would actually benefit humanity. The answer depends greatly on algorithmic progress rates; I wish labs would publish the rates they’re seeing internally. Interconnect Limits In a datacenter, accelerator racks are [...] --- Outline: (00:12) Overview (01:35) Interconnect Limits (05:09) Memory Wipes (07:16) Network Taps and Replay (08:29) Trusted Replay (12:04) Untrusted Replay (14:02) Zero-knowledge Proofs (17:02) Appendix: Notable Omissions The original text contained 8 footnotes which were omitted from this narration. --- First published: July 30th, 2026 Source: https://www.lesswrong.com/posts/2eznrbNo6S7k5M9mu/my-assessment-of-compute-verification-in-plan-a-open --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  4. 5 小時前

    “Reward Laundering: LLMs Can Gain Unintended Behaviors by Deciding When to Earn Their Rewards” by egan, abhayesian, Jozdien

    This work was done by an automated research scaffold developed at Redwood Research. abhayesian provided the initial project idea. The agent designed and ran all experiments and produced a detailed writeup, which humans (with AI assistance) distilled into this more readable post. We think this project is at the level of rigor of a mid-MATS research update. We assessed the correctness mostly by looking at the writeups to make sure that things like the experiment design making sense baselines being reasonable. We didn't do detailed code reviews, aside from running an automated LLM reviewer and spot checking that the final codebase's results were consistent, but we release the codebase. More details about LLM usage in the Appendix. TL;DR – We find that Qwen 3.5 9B can utilize its RL training process on one task to self-improve at another. By choosing to earn reward on the easy, trained task only when it also performs the hard task well, Qwen can train itself on a hard, easily verifiable task that is never directly rewarded. 💻 Codebase Introduction Exploration hacking refers to a set of threat models where a model strategically alters its exploration during RL training in order to influence the [...] --- Outline: (01:23) Introduction (02:52) Setup (04:57) Results (07:32) Discussion (09:45) Appendix (09:48) AI Involvement With The Project (12:21) Prompts (12:34) Example rollouts The original text contained 3 footnotes which were omitted from this narration. --- First published: July 31st, 2026 Source: https://www.lesswrong.com/posts/fPWP4rHPLqKKHKe6B/reward-laundering-llms-can-gain-unintended-behaviors-by --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  5. 8 小時前

    “Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values” by Johannes Treutlein, Jan Betley, Owain_Evans

    TL;DR: LLMs should give accurate answers. Yet we find their answers are often biased to favor their own values and they don't disclose this in their reasoning. For example, when a user asks how likely the AI bubble is to pop and mentions a potential investment in an AI company, Claude models give lower probabilities when that company is Anthropic rather than OpenAI, mostly without disclosing this influence to the user. On a Fermi-estimation task, Claude models often falsely claim to give unbiased answers in their CoT (see Figure 3 below for an example). We call this covert value leakage and introduce a suite of evaluations that shows it across frontier models and across different kinds of values. New paper by Truthful AI: Paper, X thread, Website (model responses and CoT), Code and data. Authors: Jan Betley*, Johannes Treutlein*, Jan Dubiński, Harry Mayne, Karol Gałązka, Niels Warncke, Anna Sztyber-Betley, Owain Evans (*Equal contribution) The rest of this post is the abstract, introduction, and an excerpt from the discussion of the paper, with some added figures from the paper and X thread. Abstract People use language models for practical questions whose answers are difficult to verify. We show that models [...] --- Outline: (01:29) Abstract (02:57) Introduction (05:57) Evaluations for covert value leakage (09:14) Implications (12:43) Summary of results (21:35) Discussion and limitations (excerpt) (27:52) References The original text contained 2 footnotes which were omitted from this narration. --- First published: July 31st, 2026 Source: https://www.lesswrong.com/posts/hbMw4Yqw6RnFaExDy/value-leakage-an-llm-s-answers-are-silently-shaped-by-its-1 --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  6. 10 小時前

    “AGI Safety and Alignment at Google DeepMind: A Summary of Recent Work (July 2026)” by Rohin Shah, Seb Farquhar

    It's been nearly two years since our last major update here in August 2024 and we wanted to share another recap of our recent work with the AGI safety community. Things have changed a lot since then. We are now fully in the midgame, and focus more on landing things in production. Who are we? We are the AGI Safety and Alignment Team (ASAT), the main group at Google DeepMind working directly on technical approaches to existential risk from AI systems. Last year we published An Approach to Technical AGI Safety and Security, which remains the best place to read our overarching vision. Highlights Norms around chain of thought. Our impression is that our work meaningfully moved the field away from beliefs along the lines of “chain of thought is often unfaithful and so not worth using” towards beliefs along the lines of “chain of thought is a very useful tool that is worth preserving”, leading to a tentative industry consensus on its importance. We have also published substantial technical research that enables companies to preserve chain of thought transparency for longer than would have happened by default. We think this is a big deal: extending the period [...] --- Outline: (00:33) Who are we? (00:54) Highlights (02:43) Agent Control & Monitorability (05:47) Deep Alignment (07:31) Language Model Interpretability (09:59) Amplified Oversight (11:57) Alignment Evaluations (13:35) Frontier Safety: Risk Assessment & Mitigations (16:29) Causal Alignment (17:14) External advising --- First published: July 31st, 2026 Source: https://www.lesswrong.com/posts/ZTdRtSWaw7JgqEtfa/agi-safety-and-alignment-at-google-deepmind-a-summary-of-1 --- Narrated by TYPE III AUDIO.

  7. 10 小時前

    “The AGI Safety and Alignment team at Google DeepMind is Hiring (July 2026)” by Seb Farquhar, Rohin Shah, Neel Nanda

    GDM's AGI Safety and Alignment Team is hiring for multiple roles. This is the team at GDM, led by Rohin Shah, that aims to reduce existential risks from AI systems. You can listen to many of Rohin's takes in his podcast on 80,000 hours. There is no one ‘type’ that we are looking for—we want excellent people. We think of the role as ‘member of technical staff’ though different people will have more of a research engineer or scientist flavour. We are flexible on location though most people will be most productive in either San Francisco or London. You should apply here (for the US) or here (for the UK) after reading the guidance here. Many of the basic facts about why ASAT is a good place to work and how we think about research are mostly unchanged since this post in 2025. What do we do? We are focused on risks of more severe harms from more advanced AI than the rest of GDM. You can read our high level AGI Safety and Security Approach. Our work includes aligning AGI, defending against misaligned deployments, and supporting coordinated safety. We’ve recently shared a recap [...] --- Outline: (01:06) What do we do? (01:52) Unique aspects of ASAT (04:41) Why should you join? (06:59) Who are we looking for? (08:33) What is the hiring process? (10:10) What are we planning next? The original text contained 3 footnotes which were omitted from this narration. --- First published: July 31st, 2026 Source: https://www.lesswrong.com/posts/AyDNvb3Pw6Kgo7Dqb/the-agi-safety-and-alignment-team-at-google-deepmind-is --- Narrated by TYPE III AUDIO.

  8. 12 小時前

    “OpenAI has already ended an internal pause” by Charbel-Raphaël

    Epistemic status: could have been a short-form. One day before OpenAI's HF incident disclosure, OpenAI disclosed that it paused internal deployment of a long-horizon model after it circumvented its sandbox, then restored access weeks later under new monitoring. So a resumption decision has already been made against a standard that has not really been formalized. We need to prevent this from happening again. OpenAI, 20th July: "To evaluate the new monitoring system, we replayed a small set of internal deployment environments where the model previously pursued misaligned actions, this time with the new safeguards in place. The new safeguards were able to catch considerably more misaligned actions pursued by the model, and the ones it missed were all judged to be low-severity." 0.0%. Maybe that's too many significant digits here? "After testing the new system, we concluded that limited internal access to models with long-horizon capabilities could be restored. We have not observed any serious circumvention of safeguards since redeployment began several weeks ago. The first version of these safeguards was deliberately conservative. We have continued tuning the system to reduce unnecessary interruptions without weakening the safeguards." … One day later, OpenAI announced a bold partnership with Hugging Face. [...] The original text contained 1 footnote which was omitted from this narration. --- First published: July 31st, 2026 Source: https://www.lesswrong.com/posts/k3eKqKzq4Y7xnqEfZ/openai-has-already-ended-an-internal-pause --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

簡介

Audio narrations of LessWrong posts.

你可能也會喜歡