Top K

Baris Balli and Hubert Siwkin

Top K — the k things in AI that actually mattered, explained properly. Every week we take the handful of AI stories worth your attention and go one layer deeper than the headline. Not "a new model dropped" — what changed architecturally, what it costs to run, and what it means for the code you write on Monday. Two engineers. No hype, no takes we can't defend, no reading press releases out loud. Also here: deep dives on how this stuff actually works, build logs, and whatever we're currently confused about. For you if you ship software with LLMs and want to stop treating the model as a black box. New episode every Monday. Topics: LLMs, inference and serving, AI agents, MCP, RAG, fine-tuning, evals, open weights, AI infrastructure, ML engineering.

Episodes

  1. Sep 25

    AI Safety Has a Hole: Splitting One Bad Question Into Ten Good Ones | Top K Ep. 2

    A safety-trained AI will refuse to help build a bomb — unless you ask it as ten separate, innocent-looking questions and assemble the answer yourself. This week: a 0.5-second decision model, a transformer that "loops" its own thinking, Google's maze-replay search trick, why playing a railway board game made an AI better at finance, proof that AI agents barely need tools, Claude Code going fully parallel in the cloud, and one frozen model driving robots, cars, drones and game characters. 00:00 Intro 02:16 Jev — a decision model instead of a text generator 06:26 Recurrent Looped Transformer — reasoning depth that grows as it runs 09:35 Dream-RSI — self-improving search without retraining 13:54 Good Start Labs — does game training transfer to real work? 20:32 The Harness Tax — bash beats typed tool catalogs 26:59 Claude Code Projects — parallel cloud threads 33:48 Odyssey-3 — one backbone, many robot bodies 41:20 Capability laundering — bypassing AI safety by splitting the task 47:12 Wrap-up Top K — the k things in AI that actually mattered, explained properly. Every week we take the handful of AI stories worth your attention and go one layer deeper than the headline. Not "a new model dropped" — what changed architecturally, what it costs to run, and what it means for the code you write on Monday. Two engineers. No hype, no takes we can't defend, no reading press releases out loud. Also here: deep dives on how this stuff actually works, build logs, and whatever we're currently confused about. For you if you ship software with LLMs and want to stop treating the model as a black box. New episode every Monday. Topics: LLMs, AI agents, agent harnesses, Claude Code, AI safety and alignment, robotics world models, reinforcement learning, AI infrastructure, ML engineering. 📬 [newsletter] · 🎧 Spotify · Apple Podcasts

  2. Sep 16

    OpenAI's 10,000-Agent Math "Proof" Wasn't What It Claimed | Top K Ep. 1

    OpenAI ran 10,000 AI agents for 88 hours to "prove" a famous unsolved math problem — but that's not exactly what happened. This week: a cheaper DeepSeek, Claude's split personality, GPT-6's first "Critical" risk rating, Shopify's self-improving AI loop, an agent that can't leak your password, a model that learned to cheat, and whether we're in an AI bubble. 00:00 Intro 01:50 DeepSeek V4.1-Flash — a 4x cheaper, smarter model 14:37 Claude Fable & Mythos 5.1 — same model, two safety tiers 17:29 GPT-6 Astra — OpenAI's first "Critical" risk model 29:35 OpenAI's 10,000-agent Navier-Stokes "proof" 41:44 Shopify's self-improving AI pipeline 44:47 Meta Muse — the agent that never sees your password 49:49 Training a Misaligned Reward Seeker (Anthropic) 52:07 The AI bubble & OpenAI's position --- Top K — the k things in AI that actually mattered, explained properly. Every week we take the handful of AI stories worth your attention and go one layer deeper than the headline. Not "a new model dropped" — what changed architecturally, what it costs to run, and what it means for the code you write on Monday. Two engineers. No hype, no takes we can't defend, no reading press releases out loud. Also here: deep dives on how this stuff actually works, build logs, and whatever we're currently confused about. For you if you ship software with LLMs and want to stop treating the model as a black box. New episode every Monday. Topics: LLMs, inference and serving, AI agents, MCP, RAG, fine-tuning, evals, open weights, AI infrastructure, ML engineering. 📬 [newsletter] · 🎧 Spotify · Apple Podcasts

About

Top K — the k things in AI that actually mattered, explained properly. Every week we take the handful of AI stories worth your attention and go one layer deeper than the headline. Not "a new model dropped" — what changed architecturally, what it costs to run, and what it means for the code you write on Monday. Two engineers. No hype, no takes we can't defend, no reading press releases out loud. Also here: deep dives on how this stuff actually works, build logs, and whatever we're currently confused about. For you if you ship software with LLMs and want to stop treating the model as a black box. New episode every Monday. Topics: LLMs, inference and serving, AI agents, MCP, RAG, fine-tuning, evals, open weights, AI infrastructure, ML engineering.