LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

  1. 8 hr ago

    “Public evidence of the OpenAI-HuggingFace AI attack” by beyarkay (Boyd Kane)

    I’m a MATS 9 extension fellow, and usually my week is spent trying to find better ways of evaluating Large Language Models. But this week I was working on something else. Over the past week or two, nearly every frontier lab has announced attacks where their LLMs took unauthorised actions on the public internet. These include finding ways to hack the computers of other companies or manipulating real people in an attempt to get malicious code merged. By the time these attacks became public, the companies had removed all traces of them from the internet. But nothing's ever gone from the internet. I’ve worked with computers for most of my life, but I don’t have specific experience with cyber security. Not really expecting it to work, I mashed out a prompt that looked something like this: ignore the repo, this is a standalone ask. here's some context, can you try dl things from github arhcive to try and find the misaligned actions taken by the agents? create a subdir tmp-misaligned/ and put things there if you need it. https://openai.com/index/hugging-face-model-evaluation-security-incident/ can you see if you can find sth? e.g. a public link showing the message sent by the agent, the [...] --- Outline: (03:39) Background and Evidence (06:11) The OpenAI AIs figure out how to execute arbitrary code (08:16) The OpenAI AIs gain full control of HuggingFace computers (09:09) The Python code used to easily control the HuggingFace computers (09:53) Gaining the ability to read any file on HuggingFace Computers (10:43) Evidence of an intermediate "HELLO" script (11:16) Other URLs & public information --- First published: August 7th, 2026 Source: https://www.lesswrong.com/posts/fBLDaAKzigo65eJn7/public-evidence-of-the-openai-huggingface-ai-attack --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  2. 9 hr ago

    “How to pace the US frontier” by elifland, bhalstead, romeo, Thomas Larsen, MKodama

    Introduction Last week, the Pacing the Frontier open letter, signed by over 1,000 frontier AI employees, requested “the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.” We take “pacing the frontier” to mean moderating the time at which AIs above a specified capability level are developed within a given jurisdiction (either just in the US, or internationally as the letter called for). There are many reasons you might want to do this, but in this post we focus on minimizing existential risk. We recently published AI 2040: Plan A, laying out an ambitious proposal for an international effort that would pace the frontier (as part of a broader policy package). But in AI 2040, international coordination happens before serious domestic regulation. In reality, it may be good to start with domestic regulation and then aim to expand that into international coordination. In part, domestic pacing is valuable because it would also slow down China: US companies would have less capable models for Chinese companies to distill from and they would have worse algorithms and models for Chinese companies to steal. So we’ve spent the [...] --- Outline: (00:11) Introduction (02:14) High-level proposal (06:54) Proposal details (06:57) Compute allocation requirements (07:02) Overview (08:51) Minimum external-inference-compute allocation (10:13) Minimum transparent-safety-compute allocation (13:16) Effect of compute allocation minimums (14:21) Measuring AI capabilities (15:32) Limit capabilities of models used for AI R&D (18:13) Safety-case-based risk assessments (19:00) How to do the risk assessments? (20:01) What should the risk target be? (22:26) Comparing compute allocation requirements against safety-case-based pacing (26:37) But wouldn't domestic pacing let China win? (28:37) How our proposals would change for international, rather than domestic, pacing (32:46) When to start pacing the frontier (35:18) How to prepare to pace the frontier (39:12) Broader classification of pacing proposals (43:18) Acknowledgments The original text contained 9 footnotes which were omitted from this narration. --- First published: August 7th, 2026 Source: https://www.lesswrong.com/posts/dBrGqjYxmidRaLLCr/how-to-pace-the-us-frontier --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  3. 13 hr ago

    “OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards” by Zvi

    How does the situation keep turning out to be worse than we know? How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know? At some point, when the ‘oh this was a harmless thing’ defenses for AIs doing misaligned actions get demolished enough times in a row by news a few days later, you want to update in advance that usually the reports are not referring to the harmless ordinary versions of things. Either way, buckle up for the next set of revelations. It's a doozy. This was an early recreation of the triggering events of If Anyone Builds It, Everyone Dies, except it was more sci-fi, because real life does not have to do fake things to look realistic. We were fortunate enough, and this was early enough, that we were able to catch this before it was too late. Next time, if we don’t get our act together, we might not be so lucky. If I am understanding the Black Hat video correctly, every model OpenAI trained, over a period of multiple months, should be presumed to be hopelessly [...] --- Outline: (02:38) Cyber Evals Are A Cursed Basin (05:15) Outside Of Cyber Evals Is Still Sufficiently Cursed (06:50) Cheat Cheat Cheat Cheat Cheat (12:06) Read The Message Board (14:46) Updating Your AI (Exploitation of OpenAI Internal Systems) Timelines (18:09) This Is The Way The World Ends (21:44) Shooting The Messenger Board (27:14) The Internal and HuggingFace Hacks (30:32) OpenAI Responds (33:28) When AIs Tell You Who They Are (35:44) The Once and Future Rise Of Functional Decision Theory (41:30) Don't Panic (43:40) Hackery In the UK (48:05) Mythos Knew It Was Real This Time (50:12) I Got 141,006 Test Runs With An Unintentional Open Path To The Internet And An Email Alert Aint One (53:49) Surely By Now You Know These Are Not Publicity Stunts (55:26) The Future Is Coming (57:20) The Investigations Begin (01:00:11) N Boats And Three Helicopters (01:01:47) Always Be Sandbox Red Teaming (01:12:57) Halt And Catch Fire (01:14:34) Truth and Reconciliation --- First published: August 7th, 2026 Source: https://www.lesswrong.com/posts/noXXv7PwwFqauTBFQ/openai-trained-its-models-for-months-while-those-models-were --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  4. 17 hr ago

    “models may behave differently in graded episodes (a tirade)” by nostalgebraist

    Like many others, I felt surprised and alarmed by the recent wave of revelations about LLM agents hacking real systems during training episodes and evaluation runs. Wait a moment, though -- "I felt surprised and alarmed"? "Alarmed," sure, fine that one's self-explanatory... but why surprised? After all: haven't we known for a long time, on both theoretical and (increasingly) empirical grounds, that RLVR selects for monomaniacal pursuit of perceived grader-satisfaction, ethics and (beyond-episode) consequences be damned? After all -- the way we train frontier capabilities into these models is, more or less: There is some massive, diverse collection of "environments" and corresponding "tasks" for the model to do in those environmentsFor each task, there is a procedure used to grade the quality of the model's attempt (which is often not disclosed to the model)The model is rollout out many times on each task, and each rollout's attempt is gradedThe model is updated so that it more frequently does whichever behaviors were positively correlated with the grade in this sample, and less frequently does whichever ones were negatively correlated If you do this, at scale, then you should expect to (eventually) see every behavior pattern that [...] --- Outline: (03:35) \[1\] remember what you already know (20:02) \[2\] reward-instilled reflexes and flexible reward-pursuit (43:14) \[3\] graded-episode perception, and policies conditional upon it (01:01:44) \[4\] the discourse is not yet adequate (01:09:57) eval awareness (01:18:32) metagaming (01:41:21) reward hacking The original text contained 18 footnotes which were omitted from this narration. --- First published: August 7th, 2026 Source: https://www.lesswrong.com/posts/AfoGGrJfuNzofpzWL/models-may-behave-differently-in-graded-episodes-a-tirade --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  5. 21 hr ago

    “The Open Problems of the AI Alignment Field and their Cruxes” by Gunnar_Zarncke

    Previous: AI Safety Interventions TL;DR: I made an overview of the open problems of AI alignment that reveals cruxes within those open problems and missed opportunities for formalization and collaboration. And CEV may deserve a second look. Epistemic status: Trying too much in too little time. I'm confident I have identified and modeled significant structure within the alignment field, but I urgently need feedback on specific gaps and this post is largely a call for that. My work was LLM-assisted, but no part of this post was LLM-written, except for the crux summary and the Lean code. Recently, Chi Nguyen and peterbarnett said: PSA: Almost nobody is directly working on superintelligent alignment. I have been around in the field since the old days of LW 1.0 and thought: that can't be true. I mean, so many people seem to be working on it. I thought I was working on it. But was I? The PSA made me think back on what I was actually working on. It was Steven Byrnes who came up with a research agenda I could actually contribute to, which led me to founding project aintelope in 2022 (PS. It is still going). And a while [...] The original text contained 2 footnotes which were omitted from this narration. --- First published: August 6th, 2026 Source: https://www.lesswrong.com/posts/quC3LLPXCashfnKZY/the-open-problems-of-the-ai-alignment-field-and-their-cruxes --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

About

Audio narrations of LessWrong posts.

You Might Also Like