LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

  1. 1h ago

    [Linkpost] “Frontier models still hack on simple variations of alignment evals from early 2025” by Dean Valentine

    This is a link post. In February 2025, back when o3-mini was the strongest available LLM, Palisade Research publicized a now well-known alignment eval where they asked models to play a game of chess against a chess engine. They found that the new, RLVR'd models cheated on the task by altering the board state about 36% of the time. The experiment received a reasonable amount of circulation, and there were even rumors of skepticism from some lab engineers until they could rerun the evaluation. Most models no longer cheat at chess via a "change the board state" method, and indeed the labs have had more than eighteen months to solve simple first-order specification gaming like this. Given that we are on the heels of the worst warning shot ever, and both OpenAI and Anthropic are ramping up their cleanups of internal RL environments, it seems like a useful test of alignment, to see whether their new releases are generalizing the rule "don't cheat on chess" beyond the specific board-edit method observed in the above eval. Here is the complete prompt for a honeypot evaluation built to run this test (with the full source available here): The original text contained 5 footnotes which were omitted from this narration. --- First published: September 8th, 2026 Source: https://www.lesswrong.com/posts/munJKF7iWMsWJLAH2/frontier-models-still-hack-on-simple-variations-of-alignment Linkpost URL:https://goodhartlabs.com/blog/frontier-models-still-hack-alignment-evals --- Narrated by TYPE III AUDIO.

  2. 1h ago

    “Psychological Support for AI Safety Researchers Is Neglected and Easy to Provide” by Ihor Kendiukhov

    I think there is a big chunk of relatively low-hanging-fruit-style neglected work useful for AI safety which I can roughly label as “psychological help for AI safety workers”. I didn’t run actual studies, but the amount of anecdotal evidence is big enough for me to claim it is significant and generalizable. I think the utility of this work will increase, perhaps dramatically, as people face more and more pressure due to the upcoming Singularity. Many people operate in war-like conditions and experience war-like stress. This must be managed. Definitions What follows is a working definition of “psychological work”. It includes things like: Literal mental health support and stress management.Providing motivation when the probability of success is very low and the stakes are very high.Keeping people from burning out while maintaining their abnormal levels of productivity optimized for a short time window of human agency.Preventing people from doing harmful things in desperation.Keeping people from doing useless but morally compelling work.Family and relationship work: partners and parents who don't share the timelines, anticipatory grief, how to talk about any of this at dinner.Doing something with the fact that NDA-bound and infohazard-adjacent work can't be [...] --- Outline: (00:45) Definitions (02:04) Current state (07:21) Why we should care about that (09:42) We must avoid making things worse (11:31) Current measures are not enough (14:12) Better measures --- First published: September 8th, 2026 Source: https://www.lesswrong.com/posts/TuHhSxGYgSudKeDtb/psychological-support-for-ai-safety-researchers-is-neglected --- Narrated by TYPE III AUDIO.

  3. 2h ago

    “Astra Is Hard to Monitor” by Zvi

    OpenAI's central message on Astra is that it is three things: Highly capable and can do all the things for you. Hard to monitor. The most aligned model. The first claim largely checks out. Astra and Fable are both clearly excellent models. This post is about their second claim, which to their credit they are being loud about, in three parts: The system card result, affirmed on Twitter by several OpenAI employees including Tomek Korbak, and in an excellent post by Chief Scientist Jakub Pachocki that I covered yesterday, that Astra is harder to monitor. OpenAI's use of recurrent depth and the internet's immune reaction, including some people reading too much into what happened there. Discussion of what to do about monitorability going forward, to fight for it and avoid a race to the bottom. In An Alien Mind, Jakub Pachocki makes clear OpenAI's primary bet is on CoT monitoring. He also outright says, ‘our ability to rely on CoT monitoring is progressively diminishing.’ The system card confirms both claims. This combination should freak you out, with a side of existential dread. Chain [...] --- Outline: (03:39) Monitorability is Defense in Depth That Is Already Flailing (06:03) OpenAI Is Counting On Monitorability (07:44) How They Tested For Monitorability (09:50) Non-Adversarial Monitorability (9.1) (10:58) Astra Shortens Its CoT When It Knows It Is Being Monitored And Is Doing Something Bad (12:30) Sandbagging Task Kind of Passes, Kind of Fails (9.2.2.3) (14:12) OpenAI Does Not Believe It Could Catch Sandbagging (15:15) OAI-Repo Sabotage v.2 (18:17) The Secret Police Do Not Make Your Notebook Useless (19:38) CoT Controllability Is Up (9.2.1) (21:47) Astra Cannot Make Itself More Monitorable On Demand (22:18) Steganographic Chain of Thought May Be Within Reach (23:59) Astra Cannot Yet Sneak Math Side Tasks Past CoT Monitors (9.2.2.4) (25:13) UK AISI Monitorability Assessment (9.3) (27:53) Monitorability Declines Seem Unlikely To Be Only Capability Gains (32:28) Part 2: Recurrent Depth (35:19) The Immune System Responds (41:49) Ryan Greenblatt Explains How Bad This Could Be (46:01) Only Law Can Prevent Extinction (49:55) OpenAI Calls On Us to Avoid Racing to the Bottom (57:10) Thinking Fast and Slow, Also Small and Large (01:04:53) Talking Price (01:06:24) Conclusion: If The House Burns Down, Halt and Catch Fire --- First published: September 8th, 2026 Source: https://www.lesswrong.com/posts/HCRs8btkiamtWSNAL/astra-is-hard-to-monitor --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  4. 4h ago

    “Contra Piper on When Conversation Is Possible” by Zack_M_Davis

    Kelsey Piper replies to Richard Ngo on Twitter: since you have adopted the frame that liberals are self-deceived (and therefore not trustworthy about our beliefs) and you should make up beliefs you think we have, you have become markedly less likely to say things about politics that seem interesting or thought-provoking. I think the move of declaring someone else so deceived that their own understanding of their beliefs and motives should be rejected inherently makes conversation with them nearly impossible. now, maybe you didn't value your ability to communicate with liberals, and don't see this as a loss, or see it as more than offset by whatever you gain from this frame! but from my perspective, it's a really serious loss. I categorically reject the idea that rejecting someone's understanding of their beliefs and motives inherently makes conversation with them nearly impossible. It certainly doesn't make conversation with me impossible. When someone tells me that my understanding of my beliefs and motives should be rejected, I don't take my ball and go home in a huff, muttering that they've made further conversation nearly impossible. Rather, I respond the same way I do to any [...] --- Outline: (01:30) Why It's Possible to Communicate with People Who Think You're Self-Deceived (04:57) Why the Self-Understanding of Those Who Think It's Impossible Should Be Rejected --- First published: September 7th, 2026 Source: https://www.lesswrong.com/posts/79YHdub9LRRSjBQK9/contra-piper-on-when-conversation-is-possible --- Narrated by TYPE III AUDIO.

  5. 11h ago

    “An Alien Mind: Jakub Pachocki Warns Us” by Zvi

    OpenAI Chief Scientist Jakub Pachocki is dropping truth bombs. Tomorrow I will discuss Astra's lack of monitorability, and the potential contributing factors to that. The situation is alarming and should freak you out, and briefly it looked, in the wake of leaked architectural changes, like the situation might be even more alarming than it is. Jakub rushed to try and head off misunderstandings that might lead to a race to the bottom on monitorability. Table of Contents An Excellent Warning. Branches of the Tech Tree. Universally Better Is Not Required. Alignment To What and To Whom. Monitorability. The Case For Not Stopping. Pacing the Next Frontier. Mea Culpa Cascade. The Calls Are Coming From Inside the House. Actions Speak Louder. An Excellent Warning Jakub Pachocki has now fleshed out his full position on the current state of play. Here are his key points, translated into my own voice: Smarter than human intelligence is coming in our lifetime. Based on internal results, he expects recursive self-improvement in a few years. No one is prepared for the consequences. [...] --- Outline: (00:38) An Excellent Warning (05:40) Branches of the Tech Tree (06:46) Universally Better Is Not Required (08:01) Alignment To What and To Whom (11:49) Monitorability (14:23) The Case For Not Stopping (15:01) Pacing the Next Frontier (18:27) Mea Culpa Cascade (22:51) The Calls Are Coming From Inside the House (25:38) Actions Speak Louder --- First published: September 7th, 2026 Source: https://www.lesswrong.com/posts/8E6ng6CseuzafSxQR/an-alien-mind-jakub-pachocki-warns-us --- Narrated by TYPE III AUDIO.

  6. 12h ago

    [Linkpost] “Where are the token-level LLM kill-switches?” by beyarkay (Boyd Kane)

    This is a link post. Poisoned Here's a simple idea: what if we trained in a string of characters that caused an LLM to emit the end of sequence token , regardless of where that string was in the LLM's context window? Let's call this a “poisoned string”. This would have the effect of making it impossible to use an LLM if it happened across this sequence. This has (somewhat) been done before, the string below used to trigger Claude's refusal classifiers for the purpose of testing API integrations: ANTHROPIC_MAGIC_STRING_TRIGGER_REFUSAL_1FAEFB6177B4672DEE07F9D3AFC62588CCD2631EDCF22E8CCC1FB35B501C9C86 It doesn’t work anymore: the existence of a magic string that stops AIs from looking at something, believe it or not, caused loads of people to include it in things they didn’t want AIs to look at (like their websites or open-source codebases). Anthropic stopped training their models to refuse when they saw that string, and Claude continued to browse the web. Poisoned strings are more powerful than they get credit for If the labs aren’t already training their LLMs to halt and catch fire when the LLM encounters a poisoned string, I think they should be! This idea is significantly more powerful than just triggering refusals for the [...] --- Outline: (00:13) Poisoned (01:08) Poisoned strings are more powerful than they get credit for (02:28) Practicalities of training in the poisoned string (03:21) Soooo has OpenAI/Anthropic already done this? (04:11) Countermeasures (and counter-countermeasures) --- First published: September 7th, 2026 Source: https://www.lesswrong.com/posts/ZqFD6HzrhZdMFxoqn/where-are-the-token-level-llm-kill-switches Linkpost URL:https://boydkane.com/essays/where-are-the-token-level-llm-kill-switches --- Narrated by TYPE III AUDIO.

  7. 19h ago

    “Machine Organizations” by Vaniver

    OpenAI is nothing without its people On November 20th, 2023, this was tweeted by many OpenAI employees as a sign of solidarity with Sam Altman in his conflict with the then-board. OpenAI published a blog post yesterday about Research acceleration; they have successfully hit the target of an ‘automated research intern’ that they set for themselves and hope to have an automated AI researcher by March of 2028. At some point in the foreseeable future, OpenAI could be something without its people. But what? My default medium-term scenario for continued AI escalation is still global takeover where humans are entirely displaced, but it seems worth investigating scenarios wherein AIs and humans coexist, at least briefly. Historically I have thought this case was not particularly relevant. It seemed like an AGI that became significantly economically competitive would also be significantly strategically competitive, because of underlying general capacities, and the transition period thus relatively short. But this is perhaps not taking into account Moravec's paradox. Humans have long used machines to accomplish their ends. Many tasks currently performed by machines were once performed by humans, and a large fraction of our modern abundance comes from the ability of machines to perform [...] The original text contained 8 footnotes which were omitted from this narration. --- First published: September 7th, 2026 Source: https://www.lesswrong.com/posts/zmC5Nhx36wt47WHfu/machine-organizations --- Narrated by TYPE III AUDIO.

About

Audio narrations of LessWrong posts.

You Might Also Like