This episode examines "Mechanist," a multi-agent system built by researchers at Zhejiang University, NUS, Southern University of Science and Technology, Heriot-Watt, UC San Diego, and Northeastern University to automate mechanistic interpretability research itself, rather than automating experiments in an external domain like chemistry or biology. The discussion covers how a central orchestrator coordinates four agents—hypothesis, experiment, verification, and iteration—drawing on a 13,000-paper interpretability knowledge graph and a 43-million-paper cross-disciplinary graph called SciAtlas to generate and test theories about how models actually compute. Key concepts explored include subliminal learning, where a trait transfers from teacher to student model through data that looks unrelated to it, and a three-frame belief decomposition (World Knowledge, Personal Belief, Attributed Belief) used to probe whether models genuinely separate fact from attributed belief. The episode previews four escalating case studies, starting with the discovery of a previously unflagged multimodal safety risk and building toward using mechanistic theories to directly intervene on model internals and even steer a biological system, raising the question of whether an AI system can meaningfully explain the black box that produced it. Sources: 1. Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence — Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen, 2026 http://arxiv.org/abs/2608.12036 2. Language models transmit behavioural traits through hidden signals in data — Alex Cloud, Minh Le, James Chua, Jan Betley, Anna Sztyber-Betley, Sören Mindermann, Jacob Hilton, Samuel Marks, Owain Evans, 2026 (Nature) https://scholar.google.com/scholar?q=Language+models+transmit+behavioural+traits+through+hidden+signals+in+data 3. Subliminal learning is a lora artifact — Todd Nief, Harvey Yiyun Fu, Mark Muchane, Ari Holtzman, 2026 https://scholar.google.com/scholar?q=Subliminal+learning+is+a+lora+artifact 4. Language models cannot reliably distinguish belief from knowledge and fact — Mirac Suzgun, Tayfun Gur, Federico Bianchi, Daniel E. Ho, Thomas Icard, Dan Jurafsky, James Zou, 2025 (Nature Machine Intelligence) https://scholar.google.com/scholar?q=Language+models+cannot+reliably+distinguish+belief+from+knowledge+and+fact 5. Genome modelling and design across all domains of life with evo 2 — Garyk Brixi, Matthew G. Durrant, Jerome Ku, et al., 2026 (Nature) https://scholar.google.com/scholar?q=Genome+modelling+and+design+across+all+domains+of+life+with+evo+2 6. Towards end-to-end automation of ai research — Chris Lu, Cong Lu, Robert Tjarko Lange, Yutaro Yamada, Shengran Hu, Jakob Foerster, David Ha, Jeff Clune, 2026 (Nature) https://scholar.google.com/scholar?q=Towards+end-to-end+automation+of+ai+research 7. Sleeper agents: Training deceptive llms that persist through safety training — Evan Hubinger et al., 2024 https://scholar.google.com/scholar?q=Sleeper+agents%3A+Training+deceptive+llms+that+persist+through+safety+training 8. Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers — Jan Dubinski, Jan Betley, Anna Sztyber-Betley, Daniel Tan, Owain Evans, 2026 https://scholar.google.com/scholar?q=Conditional+misalignment%3A+common+interventions+can+hide+emergent+misalignment+behind+contextual+triggers