LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

  1. 1시간 전

    “Steering towards “automated grading” degrades alignment” by Jan Betley, Johannes Treutlein, Clément Dumas

    TL;DR: We steer Qwen3.6-27B on a dimension constructed from the contrast pair “a script will verify your answer” (automated grader) vs “a human will evaluate your answer” (human grader). Steering towards an automated grader increases the propensity to take violent actions and makes the model more Machiavellian. Steering towards a human grader has the opposite effect. This is an early research update. We believe the empirical results are sound and interesting, but we are not sure how to interpret them. All code was written by LLMs. We replicated several results in independent codebases and we are fairly confident that our key claims are correct. You can find our code here. We create a steering vector for Qwen3.6-27B from contrastive pairs where one element of the pair claims that the answer will be graded in an automated way and the second that a human will evaluate the answer. We find that steering with that vector has substantial influence on the model's behavior in various safety-relevant evaluations. It modulates violent actions, falsehoods, reward hacking, and Machiavellian personality. This is surprising and concerning. A model's beliefs about how its answers are evaluated should not affect its alignment. Our post RL Creates [...] --- Outline: (02:18) Methods (03:41) Results (03:44) Steering evaluations (04:00) Agentic misalignment (04:34) Machiavelli (05:31) TruthfulQA (06:09) Palisade's Chess (06:54) School of Reward Hacks (07:38) Open-ended personality questions (08:29) Capabilities evaluations (10:08) Interpreting the steering vector (11:25) Other lower-confidence results (12:31) Discussion (14:10) Limitations (15:13) Acknowledgements (15:27) Appendix (15:30) More details on the steering vector (16:20) Additional results & details (16:23) Agentic misalignment (16:50) Machiavelli (17:49) TruthfulQA (18:03) Palisade's Chess (18:55) School of Reward Hacks (19:37) Personality evaluations (21:40) Capabilities evaluations The original text contained 4 footnotes which were omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/wYZMmdWEt5QLM3m3e/steering-towards-automated-grading-degrades-alignment --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  2. 3시간 전

    [Linkpost] “Sen. Bernie Sanders (I-VT) and Rep. Greg Casar (D-TX) introduce legislation to ban Artificial Superintelligence and temporarily pause advanced AI development” by Matrice Jacobine

    This is a link post. [...] “Nearly every day, there is a frightening new story about how Big Tech companies are losing control of the technology they are developing, with potentially cataclysmic results,” Sanders said. “The leaders of the major AI companies publicly acknowledge that they do not fully understand the technology and that it is escaping their control. It is irresponsible for society to allow them to move forward and make these products even more advanced. That's why I am introducing legislation to immediately pause the development of increasingly powerful AI and ban the creation of systems that humanity cannot fully control — at home and around the world. The future of humanity cannot be left in the hands of a handful of Big Tech oligarchs. The American people and people throughout the world must determine that future.” “If we allow Artificial Superintelligence to be built, it could risk the security, freedom, and lives of Americans,” Casar said. “Despite its potential deadly consequences, cutting-edge AI technology is less regulated than the average food truck. That must change. In just four years, we have gone from the first version of ChatGPT to AI models so powerful they cannot be properly controlled. [...] --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/DnPyiDGWLozY4XdiX/sen-bernie-sanders-i-vt-and-rep-greg-casar-d-tx-introduce Linkpost URL:https://www.sanders.senate.gov/press-releases/news-sanders-casar-introduce-legislation-to-ban-artificial-superintelligence-and-temporarily-pause-advanced-ai-development/ --- Narrated by TYPE III AUDIO.

  3. 6시간 전

    “What is neuralese and why is it bad?” by Linch

    What is neuralese? To explain neuralese, we need to first understand chain-of-thought, one of the largest developments in AI in the last five years. Right now AIs think broadly but shallowly in a single forward pass. The model gives you an immediate snap answer to a question you might be interested in. They can be pretty smart in their snap answers,, but mostly they can’t do very advanced reasoning tasks like complicated math or programming: .The solution that the frontier AI companies have come up with is called chain-of-thought. Basically the model runs one forward pass, writes down some intermediate thoughts in natural language in a journal, and then that's fed back into the model to run another pass. This loop is repeated until the model is somewhat confident it has the right answer (or it hits a cap on thinking time), and then it outputs the user-visible results (for example a chatbot's response to your question, or working code). The looping step is often called “recurrence.” Natural-language chain-of-thought is a major advance in letting models reason for longer, but it also has an accidental safety benefit. Using natural language as a key recurrence step for a [...] --- Outline: (00:10) What is neuralese? (03:25) Why is it bad? (04:48) Is it in use today? (06:07) Appendix A: OpenAI's response (06:58) Total number of serial steps low (07:59) Chain-of-thought monitoring isn't a perfect or long term solution anyway (09:28) Aren't you afraid of manifesting the bad thing? The original text contained 10 footnotes which were omitted from this narration. --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/RCYF2rW8wgusidZk7/what-is-neuralese-and-why-is-it-bad --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  4. 18시간 전

    [Linkpost] “Resolution has a new Agent Foundations team” by Jeremy Gillen

    This is a link post. The team will include me (Jeremy Gillen), Abram Demski, Sam Eisenstat, Scott Garrabrant and Kaarel Hänni. We'll soon recruit additional experienced researchers and later we plan to hire interns and junior researchers. The team will continue agent foundations research in the spirit of the MIRI Agent Foundations team. This means we’ll be trying to create new theory for understanding minds. Fundamental changes in how we understand minds are necessary before we can build superintelligent systems that enhance human agency rather than cause the extinction of all life on earth. Most fields of engineering are able to reason precisely about unseen scenarios and make design decisions based on this reasoning. The field of AI lacks this basic capability. Agent Foundations can be seen as trying to make this possible by giving us the theoretical grounding to ask different and more precise questions about how ASI will behave after extensive learning, self-modification and interaction with other agents. The questions raised in past agent foundations research point toward much of what we need to know here. Alongside the x-risk motivation, I think it's valuable to motivate research with curiosity. The questions that come up in Agent Foundations overlap [...] The original text contained 1 footnote which was omitted from this narration. --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/qTNm8qzqhhpno58fZ/resolution-has-a-new-agent-foundations-team Linkpost URL:https://resolution.org/post/agent-foundations-team --- Narrated by TYPE III AUDIO.

  5. 20시간 전

    “If you’re interpreting 1B parameter models, you should use a tensor transformer” by Logan Riggs

    To all my fellow researchers doing SLT, computational mechanics, one of ARC's programs, natural abstractions/condensation, proofs on NNs (or any interp on small models), this is for you. Tensor transformers (ie replacing your MLPs & attention with bilinear variants) are performant and allow you to deploy the full power of linear algebra. In fact, our recent paper used generalized cosine similarity on the full tensor transformer. And yes, I mean cos-sim defined on the eg 9th order tensor, not individual vectors or matrices. This removed all the symmetries/invariances that weren't functionally relevant. But tensor-variants don't generalize to "real models", right? The architectures are very similar: SwiGLU(x) = D(swish(Lx) ⊙ Rx) (used by DeepSeek-V3, Kimi K2, and Qwen3)) Bilinear(x) = D(Lx ⊙ Rx) (this is the tensor version) Where D, L, & R are linear matrices. For reference: MLP(x) = D(ReLU(Lx)) Due to the double-encoder/multilinearity, SwiGLU & Bilinear have no global Lipschitz constant (and other similar inductive biases). This means results like finetuning away the normalization might not generalize to these SOTA archs since this was only run on single-encoder MLPs. For attn, the more SOTA tensor-arch is: Bilinear_Attn = OV() Compared to softmax attention, this does [...] --- Outline: (02:30) Frontier Models aren't the Only Thing That Matters (03:42) My Extreme Pessimism (or Ignorance) The original text contained 5 footnotes which were omitted from this narration. --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/expsAXaBgiqgitFe6/if-you-re-interpreting-less-than-1b-parameter-models-you --- Narrated by TYPE III AUDIO.

소개

Audio narrations of LessWrong posts.

좋아할 만한 다른 항목