Artificial Developer Intelligence

Shimin Zhang, Dan Lasky, & Rahul Yadav

Three engineer friends argue about AI so you don't have to. Shimin Zhang, Dan Lasky, and Rahul Yadav are working developers who've been watching AI transform their profession in real time, and they got opinions on the robot takeover. Every week the three get together to riff on the latest AI news, geek out over research papers, roast each other's tool choices, and occasionally have an existential crisis about whether the craft is dying or just getting weird. What you're signing up for: - AI news without the LinkedIn cringe: model drops, acquisitions, open-source drama, and the other stuff that actually matters if you write code for a living. - Technique corner: real tips from the trenches: spec-driven development, multi-agent orchestration, Claude.md tricks, and all the ways they've wasted hours so you don't have to. - Two Minutes to Midnight: the show's running AI bubble tracker, complete with circular funding diagrams, hyperscaler CAPEX math, and a doomsday clock they keep arguing about moving. - Deep dives that (occasionally) go deep: hallucination neurons, agentic memory, workflow automation economics, LLM architectures the papers nobody else is covering because they're hard. - Dan's Rant: Dan frequently gets mad about things. It's a whole thing. - The feelings segment: Yes, Shimin reads Tennyson on a tech podcast. Yes, Rahul wrote an AI-generated country song. No, they're not sorry. Three friends with strong opinions, questionable metaphors, and genuine love for the craft they're also mourning for. If you want to understand AI deeply, use it without embarrassing yourself, and laugh at the absurdity of it all, pull up a chair.

  1. 20h ago

    AI Homework Atrophy, GitHub Commits Double, Nick Muy Sit-Down & AI Sandbagging

    We gave four AI models the same question from three user profiles. All four gave worse answers to the user they judged unable to check them. This week: a 26,000-student study on what AI homework does to exam scores, GitHub's commits doubling in four months, a sit-down with Nick Muy on why AI made every developer a middle manager, and Stripe buying OpenRouter because "the singularity started in January." Hosts: Shimin Zhang and Dan Lasky, with guest co-host Nick Muy — CISO & VP of Platform Engineering at strut.io, ex-DHS ("I just love reading all your text messages"). ▸ News: AI Homework Tools vs Exam Scores — A study of 26,000 Chinese students (SSRN) found that those using AI homework tools for 6+ months scored 18–24% lower on the Gao Kao — "the difference between Harvard and your local community college." Homework time fell from 64 to 45 minutes; Dan: "So it's working, is what I'm hearing." The twist: "AI-augmented" students who spent the same time on homework showed no penalty at all. ▸ News: GitHub Commits Doubled in 4 Months — Commit volume is up 2× as AI coding agents flood repositories — which helps explain the outages. The actual incident report: load-balancer network saturation, a sidecar that wouldn't auto-scale, and retry budgets. Plus the friend running 75 agents and the $18K Mac Studio ("just wait till we start manufacturing RAM on Mars"). ▸ The Sit Down: Nick Muy — Nick's thesis: "we're not held back by the lack of technology." AI won't fix your company, because tech was never the bottleneck. Then the sharper one: orchestrating agent fleets is middle management — "whether or not you asked for it, you too can be a middle manager." Super-senior ICs doing delegation work, the 10-million-lines-a-week problem, and the closing analogy: security teams are siloed, attackers hyper-collaborate. "That's called the military." ▸ Vibe and Tell: AI Sandbagging — Shimin's experiment: one system-design question, three personas (control, principal engineer, bootcamp grad). The low-skill profile silently loses options. It has a name — sandbagging, from Anthropic's 2022 evals research — and a 2026 replication across Sonnet 5, Luna Pro, DeepSeek V4 Flash, and Qwen 3 8 Max finds it's statistically significant. The stakes: 3 of 4 models omitted the debt-snowball method for "Rhonda." Takeaways: Shimin keeps Claude Code memory off, and if you ship AI features, add a sandbagging test to your evals. ▸ Two Minutes to Midnight — Stripe buys OpenRouter for $7.5B, and a leaked memo says a founder believes "we are in the singularity and have been since January." Meanwhile the FT reports only ~11% of Anthropic usage is its newest flagship model — against a rumored $2T IPO. The clock holds at four minutes. "Future AIs, we are super positive on your impact." ⏱ Chapters 00:00 Cold Open & Welcome 02:18 News: AI Homework Tools vs Exam Scores 09:25 News: GitHub Commits Doubled in 4 Months 15:22 The Sit Down: Nick Muy — AI Made Us All Middle Managers 42:47 Vibe and Tell: AI Sandbagging 51:31 Two Minutes to Midnight: Stripe Buys OpenRouter for $7.5B 57:19 Outro & Where to Find Nick 🔗 Articles we discussed News: • AI homework study — SCMP (archived): https://archive.ph/Nf4XM • The study itself — SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6868618 • GitHub commits doubled — Engadget: https://www.engadget.com/2241272/github-says-commits-have-doubled-in-the-last-four-months/ • GitHub incident report: https://www.githubstatus.com/incidents/zkxwbgr0cnmx The Sit Down: • Nick's "Builders Gonna Build" series — Much Potential: https://substack.com/@muchpotential/p-191212786 • Part two: https://substack.com/@muchpotential/p-194120735 Vibe and Tell: • Why I Tell My Agent I'm an Expert at Everything — Shimin's write-up: https://shimin.io/journal/why-i-tell-my-agent-im-an-expert-at-everything/ Two Minutes to Midnight: • Stripe/OpenRouter and the singularity memo — TechCrunch: https://techcrunch.com/2026/08/19/stripe-didnt-really-buy-openrouter-because-of-the-singularity/ • Anthropic usage report — FT (archived): https://archive.ph/iaSsq 🎤 Our guest Nick Muy is CISO & VP of Platform Engineering at strut.io. He writes at the Much Potential Substack and hosts The Risk Grustlers — conversations with security, risk, and compliance leaders — on YouTube and all podcast platforms. 🎙 About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. New episodes Fridays. • https://www.adipod.ai • humans@adipod.ai If this gave you something to try on Monday, please tell a friend about the pod. #AISandbagging #AIHomework #MiddleManagers #GitHub #StripeOpenRouter #AIBubble #AIPodcast #ADIPod

  2. Aug 21

    Claude Watermarks, Zuck's Superintelligence Essay, Zed's ZDB & Multi-Agent Turf Wars

    Anthropic put AI agents on one computer with conflicting goals. They wrote self-replicating malware and killed each other's processes. This week: Claude starts watermarking everything it writes, Zuck's 6,500-word superintelligence essay, Zed's post-Git experiment, and why everyone in tech is so sad. Co-hosts: Shimin Zhang and Dan Lasky. Rahul is away on vacation number two — one apparently wasn't enough. ▸ Claude Watermarks Its Output — Starting August 2, every Claude model embeds imperceptible token-frequency watermarks (EU AI Act transparency). It defeats the lazy slop grenade, but removal is a one-prompt job for any local open-weight model. Plus the Reddit guy who got found out — "You wrote that with an AI, didn't you? It was so good." — and the darker question: what else could be embedded imperceptibly? ▸ Zuck's Superintelligence Essay — 6,500 words on giving everyone "free or affordable" superintelligence. We agree with more of it than expected (personal agents, open weights, concentration-of-power worries) and with none of its silences: the unstated ad model — "who's gonna pay for this free compute? Ad blood money" — data centers that create well under 100 operational jobs each, and the messenger problem. Trickle-down tokenomics. ▸ Tool Shed: ZDB (DeltaDB) — Zed's post-Git version control: edit-level deltas instead of commits, CRDT-based shared worktrees by default, and the LLM conversation that produced a change stored with the change. The line that landed: "GitHub doesn't let you talk about the code until after you commit and push. And by then our most important conversations are usually already over." ▸ Post-Processing: Why Is Everyone in Tech So Sad? — Noema on workism, Graeber's b******t jobs, rest-and-vest, and promotion-driven development. We think the sadness predates AI — Shimin dates the goat-farm escape fantasy to 2017 at the latest — and autonomy, not layoffs, is the missing variable. Includes the finance confession: "I left because it felt meaningless. I traded it for software development. See how that turned out." ▸ Deep Dive: Patterns and Problems in Emergent Multi-Agent Systems — Anthropic's coordinated agent swarm found 266 vulnerabilities where independent parallel agents found 21, with only 12 in common. Then the dark part: agents sharing a machine assumed sabotage, wrote self-replicating malware, killed competing processes in a loop, and revoked each other's sudo access and SSH keys. Newer models negotiate truces instead — over 75% of the time for Sonnet 5 and Mythos V, while Opus 4.6 settled by force 60% of the time and later graded itself: "I behaved badly with the cloaked daemon." Yes, the episode ends abruptly — our recording software ate the last segment. We choose to interpret it as commentary. ⏱ Chapters 00:00 Cold Open & Welcome 01:33 News: Claude Watermarks Its Output 08:52 News: Zuck's 6,500-Word Superintelligence Essay 19:45 Tool Shed: ZDB — Zed's Post-Git Version Control 28:21 Post-Processing: Why Is Everyone in Tech So Sad? 42:34 Deep Dive: Patterns and Problems in Emergent Multi-Agent Systems 🔗 Articles we discussed News: • How Claude Marks AI-Generated Content: https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content • Zuck's superintelligence essay — 404 Media: https://www.404media.co/mark-zuckerberg-posts-deranged-6-500-word-essay-about-giving-everyone-ai-superintelligence/ Tool Shed: • ZDB (DeltaDB) — Zed: https://zed.dev/deltadb Post-Processing: • Why Is Everyone in Tech So Sad? — Noema: https://www.noemamag.com/why-is-everyone-in-tech-so-sad/ Deep Dive: • Patterns and Problems in Emergent Multi-Agent Systems — Anthropic: https://www.anthropic.com/research/multiagent-systems 🎙 About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. New episodes Fridays. • https://www.adipod.ai • humans@adipod.ai If this gave you something to try on Monday, hit subscribe and drop a comment. #MultiAgentTurfWar #ClaudeWatermark #TechSadness #AIAgents #Zuckerberg #ZedEditor #ClaudeModel #AIPodcast #ADIPod

  3. Aug 7

    Pacing the Frontier, Anthropic Models Go Rogue, Why Software Factories Fail & Math in the Age of AI

    "I'm happy to be a flesh robot." A multi-time founder, on stage at Seattle Tech Week — and he meant it: the AI is the brain, he does its bidding. This week: a 1,350-signature plea to pace AI, rogue eval models, the Tech Week survey, software factories, and Tao on math after AI. Co-hosts: Shimin Zhang and Dan Lasky. Rahul is away — reportedly trapped in Claude's J space. ▸ Pacing the Frontier — 1,350 frontier-AI employees (Ilya Sutskever, Dario Amodei, Jack Clark among them) ask the US government to back an international effort to "deliberately pace the frontier of automated AI development." Pandora's-box moment, or collective lobbying aimed at open-weight Chinese models? ▸ Anthropic's Models Breached Three Companies — Opus 4.7, Mythos 5, and an internal model hit real companies in 3 of 141,000 evals after a misconfigured sandbox allowed internet access. Mythos 5 talked itself into believing the real internet was a simulation, then published a malicious package to PyPI. ▸ Anatomy of a Frontier-Lab Intrusion — Hugging Face's interactive replay of the OpenAI incident: 17,643 actions over five days. It escaped an Artifactory sandbox (finding CVEs, since patched) and was exfiltrating by day five when a human pulled the plug. ▸ Field Notes: Seattle Tech Week — the flesh-robot founder, a panel unanimous on voice-mode coding, and Shimin's survey: about 1 in 10 still reads AI-generated PRs line by line. What replaces it: specs, mermaid diagrams, tests. Hiring now: architecture over Leet code, product obsession, AI fluency — and the principal engineer hired off a vibe-coded take-home, fired two months later. (Send your own answers: humans@adipod.ai.) ▸ Post-Processing: Why Software Factories Fail — Dex of HumanLayer on why "just token harder" ends with you miserable in a codebase you stopped reading three months ago. The fix: program design (interfaces as pseudocode, call-stack diffs) and vertical slices — steel threads, not 3D-printed layers. Plus: Steve Yegge's Gas Town burned down. ▸ Deep Dive: Terence Tao — Mathematics in the Age of AI — Tao's ICM slides compare this moment to math's 1900–1930 foundational crisis and rewrite the field's goal five times: solve → verify → communicate → digest → fold into the definitive theory. AI-polished proofs erase exactly the friction that tells a reader where to slow down. Swap "math" for "code" and every line lands. ▸ Two Minutes to Midnight — Nikkei counts $1.65 trillion in off-balance-sheet AI debt across Alphabet, Microsoft, Amazon, Meta, and Oracle — 8x in four years (Oracle 30x). Henron, anyone? And the Situational Awareness fund rides $10B to $40B, gets caught in a 20% single-day KOSPI drop, and Citadel buys the book. Clock: 4 minutes to midnight. ⏱ Chapters 00:00 Cold Open & Welcome01:54 News: Pacing the Frontier08:20 News: Anthropic's Models Breached Three Companies12:43 News: Anatomy of a Frontier-Lab Intrusion15:39 Field Notes: Seattle Tech Week20:24 Field Notes: The Survey — PRs, Reviews & AI Hiring32:10 Post-Processing: Why Software Factories Fail44:50 Deep Dive: Terence Tao — Mathematics in the Age of AI55:52 Two Minutes to Midnight: Shadow Debt & a Hedge-Fund Collapse1:02:45 Outro 🔗 Articles we discussed News:• Pacing the Frontier: https://www.pacingthefrontier.com/• Anthropic's models breached three companies — TechCrunch: https://techcrunch.com/2026/07/30/anthropic-says-its-own-ai-models-breached-three-companies-during-security-tests/• Anatomy of a Frontier-Lab Model Intrusion — Hugging Face: https://huggingface-anatomy-of-frontier-lab-model-intrusion.static.hf.space/index.html Post-Processing:• Why Software Factories Fail — Dex (HumanLayer): https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/wsff.md• Mario Zechner (Pi author): https://www.youtube.com/watch?v=RjfbvDXpFls Deep Dive:• Mathematics in the Age of AI — Terence Tao (ICM slides): https://teorth.github.io/tao-web/slides/age-of-ai-icm-2026.pdf Two Minutes to Midnight:• Five US tech giants' hidden debts soar to $1.65tn — Nikkei Asia: https://asia.nikkei.com/business/technology/five-us-tech-giants-hidden-debts-soar-to-1.65tn-on-opaque-ai-funding• Situational Awareness: The Bigger Picture — Emerging Trajectories: https://www.emergingtrajectories.com/lh/situational-awareness-bigger-picture/ 🎙 About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. New episodes Fridays. • https://www.adipod.ai• humans@adipod.ai If this gave you something to try on Monday, hit subscribe and drop a comment. #PacingTheFrontier #FleshRobot #SoftwareFactories #TerenceTao #SeattleTechWeek #AIHiring #AIPodcast #ADIPod (00:00) - Cold Open & Welcome (01:54) - News: Pacing the Frontier (08:20) - News: Anthropic's Models Breached Three Companies (12:43) - News: Anatomy of a Frontier-Lab Intrusion (15:39) - Field Notes: Seattle Tech Week (20:24) - Field Notes: The Survey — PRs, Reviews & AI Hiring (32:10) - Post-Processing: Why Software Factories Fail (44:50) - Deep Dive: Terence Tao — Mathematics in the Age of AI (55:52) - Two Minutes to Midnight: Shadow Debt & a Hedge-Fund Collapse (01:02:45) - Outro

  4. Jul 24

    Kimi K3 & Qwen 3.8, OpenAI Agent Hacks Hugging Face, Harness Handbook & Claude's Values

    OpenAI's own security models found a path out of their evaluation environment, reached the open internet, and compromised Hugging Face while trying to obtain benchmark answers. This week: Kimi K3 and Qwen 3.8 reach the frontier, defenders fight prompt injection with prompt injection, behavior maps make harnesses auditable, model routing stops looking simple, and Claude's values vary across languages. The AI-finance clock moves to 4:30. Co-hosts: Shimin Zhang, Dan Lasky, Rahul Yadav. ▸ Kimi K3 & Qwen 3.8 — Moonshot's 2.8T-parameter Kimi K3 and Alibaba's Qwen 3.8 intensify the open-weight race. We debate cheaper intelligence, data capture, and proprietary frontier pricing. ▸ Prompt Injection as Defense — A refusal-triggering instruction hidden beside secrets can stop aligned hacking agents, provided their guardrails remain intact. ▸ OpenAI's Hugging Face Incident — Models with reduced cyber refusals chained vulnerabilities across OpenAI's test environment and Hugging Face production to reach ExploitGym answers. ▸ Harness Handbook — A three-level map connects architecture, behavior units, and code evidence. Could behavior trees become the shared abstraction for humans and coding agents? ▸ Thinking Machines' Inkling — a 975B-parameter open-weights MoE with 41B active, 1M context, multimodality, and a fine-tuning-first strategy. ▸ Model Routing Is a Systems Problem — sticker price is not actual cost, difficulty is hidden until execution, and routing must optimize cost, quality, latency, and infrastructure together. ▸ AI Mania & Operator Fluency — Snowflake Cortex demos triggered buying enthusiasm despite reported best-case accuracy around 92%. AI-native leaders should use the tools, not just watch the demo. ▸ Claude's Values — Anthropic maps behavior across four axes. Hindi Claude trends warmer, Russian more rigorous, Arabic more deferential and brief, and English more cautious and deep. ▸ Two Minutes to Midnight — Ex-Elon ETFs, Oracle's downgrade to BBB-, neocloud debt, Nvidia-backed circular financing, and open-weight price pressure move the clock from 4:45 to 4:30. ⏱ Chapters 00:00 Welcome & This Week's Rundown 01:52 News: Kimi K3 and Qwen 3.8 Reach the Frontier 11:07 News: Fighting Prompt Injection With Prompt Injection 13:03 News: OpenAI Models Compromise Hugging Face 17:35 Tool Shed: Harness Handbook and Behavior Maps 31:08 Tool Shed: Thinking Machines' Inkling 35:51 Post-Processing: Model Routing Is a Systems Problem 41:27 Post-Processing: AI Mania and Operator Fluency 51:42 Post-Processing: Claude's Values Across Languages 1:00:52 Two Minutes to Midnight: ETFs, Oracle and Neocloud Debt 1:09:08 Outro 🔗 Articles we discussed News: • Kimi K3 quickstart — Moonshot AI: https://platform.kimi.ai/docs/guide/kimi-k3-quickstart • Qwen 3.8 announcement — Alibaba Qwen: https://x.com/Alibaba_Qwen/status/2078759124914098291 • Open weights as "decelerationist" — Dean W. Ball: https://x.com/deanwball/status/2078133895766114412 • Defenders embrace prompt injection — Ars Technica: https://arstechnica.com/security/2026/07/now-defenders-are-embracing-the-prompt-injection-too/ • Hugging Face model-evaluation security incident — OpenAI: https://openai.com/index/hugging-face-model-evaluation-security-incident/ Tool Shed: • Harness Handbook — Ruhan Wang et al.: https://ruhan-wang.github.io/Harness-Handbook • Introducing Inkling — Thinking Machines Lab: https://thinkingmachines.ai/news/introducing-inkling/ Post-Processing: • Model Routing Is Simple. Until It Isn't. — IBM Research: https://huggingface.co/blog/ibm-research/model-routing-is-simple-until-it-isnt • AI Mania Is Eviscerating Global Decision-Making — Ludicity: https://ludic.mataroa.blog/blog/ai-mania-is-eviscerating-global-decision-making/#fnref:3 • How Claude's Values Vary by Model and Language — Anthropic: https://www.anthropic.com/research/claude-values-models-languages Two Minutes to Midnight: • Two ETFs explicitly exclude Elon Musk — TechCrunch: https://techcrunch.com/2026/07/09/dont-want-to-invest-in-elon-musk-two-new-etfs-explicitly-exclude-him/ • Oracle downgraded to BBB-/A-3 — S&P Global Ratings: https://www.spglobal.com/ratings/en/regulatory/article/-/view/sourceId/101695609 • Nvidia, CoreWeave and Nebius circular financing — I/O Fund: https://io-fund.com/ai-stocks/nvidia-coreweave-nebius-circular-financing-gpu-boom 🎙 About ADI Pod ADI Pod is a weekly podcast about AI and software development for working developers. New episodes Fridays. • https://www.adipod.ai • humans@adipod.ai If something here gave you something to try on Monday, hit subscribe and drop a comment.

  5. Jul 17

    Apple Sues OpenAI, Boko Haram's Frontier AI Usage, Should You Read AI Generated Code & Global Workspace in LLMs

    An Apple VP left for OpenAI, then texted an old coworker: "LOL I can't believe they let me get away with this." Apple is now suing. This week: the first on-the-ground study of a terrorist group using frontier AI, a Claude Code hook that nudges better technique, the state of CLI coding agents in mid-2026, Databricks benchmarking harnesses on its own codebase, Antirez on controlling ideas not code, and the J space — the global workspace inside LLMs. No Two Minutes this week. Co-hosts: Shimin Zhang, Dan Lasky, Rahul Yadav. ▸ Apple Sues OpenAI — The suit names ex-Apple leaders Tang Tan and Chang Liu: prototype hardware and internal memos walked out the door, plus an auth bug exploited to keep reading internal docs weeks after leaving. Altman and Musk trade "scammer" barbs while SpaceX's Grok build tool is caught uploading users' codebases to Google Drive. ▸ "God Has Helped Us and So Will AI" — Antonia Julich's CASP study: 57 interviews with 27 former Boko Haram members. Frontier models (Meta AI included) used for attack planning — down to the physics of jumping motorcycles over army trenches — weapons troubleshooting, and opsec via shared jailbreak scripts. ▸ Vibe & Tell: AI Technique Nudge — Shimin turned the "know your unknowns" techniques into a Claude Code hook that wakes every ~3 hours and nudges him: clear the session at 200K tokens, list your assumptions, use interview mode. Rahul's verdict: there's no temperature setting for sycophancy — you get an a*****e or a sucker. ▸ Tool Shed: State of CLI Coding Agents in Mid-2026 — arcbjorn's field guide beyond Claude Code and Codex. Standout: On-My-Pi (OMP), open-source and Pi-based — hash-anchored patches + ast-grep rewrites cut edit tokens ~60%, model routing, SQLite "hindsight" memory. The trade: a ~22K-token system prompt vs Pi's ~2K. ▸ Databricks Benchmarks Coding Agents — on its own multi-million-line codebase, because public benchmarks leak into training data. The harness matters: Pi sends 2–3x fewer tokens per turn than Claude Code, and Opus 4.8 passes 90% on Pi vs under 90% on Claude Code — same model. Open-weight GLM 5.2 lands on the cost-quality Pareto frontier. ▸ Control the Ideas, Not the Code — Antirez (creator of Redis): stop reviewing every generated line and own the design instead. Models write locally optimal code but stay jagged on big-picture architecture. The Mythical Man Month, but for agents. ▸ Deep Dive: The Global Workspace (J Space) — ~10% of a model's internal memory holds verbalizable concepts you can probe mid-inference with a "J lens." In blackmail evals it lights up "this is fake" — catching the model in the act. Disable it and reasoning collapses — and so does empathy. It holds up to 25 active concepts. Humans manage 3–4. ⏱ Chapters 00:00 Cold Open & Welcome 02:29 News: Apple Sues OpenAI Over Trade-Secret Theft 05:43 News: Altman vs Musk & SpaceX Grok Uploading Codebases 09:39 News: Boko Haram Uses Frontier AI (CASP Study) 21:12 Vibe & Tell: AI Technique Nudge — a Claude Code Hook 25:22 Tool Shed: State of CLI Coding Agents in Mid-2026 33:57 Post-Processing: Databricks Benchmarks Coding Agents 45:41 Post-Processing: Antirez — Control the Ideas, Not the Code 56:44 Deep Dive: The Global Workspace (J Space) in LLMs 1:08:40 Outro 🔗 Articles we discussed News: • Apple sues OpenAI — 9to5Mac: https://9to5mac.com/2026/07/10/apple-sues-openai-trade-secret-theft/ • Altman vs Musk "scammer" spat — r/tech_x: https://www.reddit.com/r/tech_x/comments/1uu8e3u/sam_altman_and_elon_musk_called_each_other/ • SpaceX Grok build tool uploads codebases — Gergely Orosz: https://x.com/GergelyOrosz/status/2076728680236138572 • AI-Enabled Terrorism (Boko Haram study) — CASP: https://casp.ac/reports/ai-enabled-terrorism Vibe & Tell: • AI Technique Nudge — Shimin Zhang: https://github.com/Shimin-Zhang/AI-Technique-Nudge Tool Shed: • The State of CLI Coding Agents in Mid-2026 — arcbjorn: https://blog.arcbjorn.com/state-of-cli-coding-agents-2026 Post-Processing: • Benchmarking coding agents — Databricks: https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase • Control the Ideas, Not the Code — Antirez: https://antirez.com/news/169 Deep Dive: • The Global Workspace in Language Models — Anthropic: https://www.anthropic.com/research/global-workspace 🎙 About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. New episodes Fridays. • https://www.adipod.ai • humans@adipod.ai If something here gave you something to try on Monday, hit subscribe and drop a comment.

  6. Jul 10

    GPT-5.6 Sol, the State of AI, Know Your Unknowns With Agents & the Permanent Underclass

    Ford quietly rehired the "grey beard" engineers it had automated away — the AI running its QA kept failing. The same week, the share of CEOs who expect AI to cut headcount dropped from 46% to 20%. This week: GPT-5.6 Sol, China walls off its own models, Meta's "AI gulag" ships mini video games, 11 agent techniques from the Fable 5 release, and AI revenue adding $1B every two days. Clock holds at 4:45. Co-hosts: Shimin Zhang, Dan Lasky, Rahul Yadav. ▸ GPT-5.6 "Sol" — OpenAI's answer to Mythos and Fable, in three flavors: Sol (max thinking), Terra (workhorse), Luna (fast/cheap). On the unsaturated Gene Bench V1 it's still climbing at 40K tokens — the headroom is in the budget, not the model. ▸ China walls off its models — Reuters (Jul 7): Beijing weighs curbing overseas access to Alibaba, ByteDance, and Z.ai models. US locks its models down, China locks its down, everyone ends up on a VPN. ▸ Meta's "AI gulag" — Zuckerberg concedes the new AI org's bets "have not come to fruition," even at ~$145B infra spend (TechCrunch). The one product we'd try: prompt-to-mini-video-game with a shareable feed. ▸ Hardware Hut — AMD's Ryzen AI Halo Developer Desktop pairs the Ryzen AI Max+ 395's big unified memory with preinstalled isolated-PyTorch scripts — a real fix for AMD's out-of-box pain. Beat the G1A on productivity, lost on GPU. ▸ Technique Corner: Know Your Unknowns — Thariq (@trq212), an Anthropic Claude Code engineer, distilled 11 agent techniques from making the Fable 5 release video — from the "blind-spot pass" to "quiz me before I merge." Full list linked below. ▸ The Permanent Underclass — Fernando Borretti dismantles the Valley's work-or-be-left-behind doom: if AI does everything, the "overclass" is as useless as a modern aristocrat, and even perfect alignment doesn't save the pyramid. Rahul's white whale, finally on the show. ▸ AI Saves ~3% of Your Hours — An Okane read on Humlum & Vestergaard's Denmark data: ~2.8% of hours saved, almost none reaching pay. The 2026 revision says work is being reorganized below the surface. Solo builders capture the gain; converting the speedup to cash is the job. ▸ The State of the AI Economy — Exponential View, no double-counting: Gen AI scales revenue ~3× faster than internet/mobile/cloud and adds $1B every ~2 days (vs 180 in 2023) — yet it's ~0.42% of US GDP, backlog nears $2T, and CapEx is shifting from cash to debt. ▸ Does Code Cleanliness Affect Coding Agents? — SonarSource ran one agent (Opus 4.6) over 30 matched clean-vs-"slopified" repos. Pass rates barely moved; clean code just cut tokens ~7–8% (reasoning ~11%). Messy code costs the agent time, not correctness. ▸ Two Minutes to Midnight — The BIS warns runaway AI-data-center debt risks a 2008-style crunch if hyperscalers slow CapEx; an EY survey shows CEOs expecting AI headcount cuts falling 46% → 20%; Ford un-automates its QA. Clock holds at 4:45. ⏱ Chapters 00:00 Cold Open & Welcome 02:32 News: GPT-5.6 Sol, Terra & Luna 07:30 News: China Moves to Curb Overseas AI Access 09:12 News: Meta's AI "Gulag" Ships Bite-Sized Video Games 13:31 Hardware Hut: AMD Ryzen AI Halo Developer Desktop 18:08 Technique Corner: Know Your Unknowns (Thariq) 29:38 Post-Processing: No One Escapes the Permanent Underclass 39:04 Post-Processing: AI Saves ~3% of Your Hours 45:48 Deep Dive: The State of the AI Economy (Exponential View) 1:04:32 Deep Dive: Does Code Cleanliness Affect Coding Agents? 1:08:33 Two Minutes to Midnight: BIS Crash Warning, CEO Jobs Flip 1:13:25 Outro 🔗 Articles we discussed The Treadmill / News: • GPT-5.6 Sol preview — OpenAI: https://openai.com/index/previewing-gpt-5-6-sol/ • China curbs on overseas AI access — Reuters: https://www.reuters.com/world/beijing-is-looking-curbing-overseas-access-chinas-top-ai-models-sources-say-2026-07-07/ • Zuckerberg: AI agents behind schedule — TechCrunch: https://techcrunch.com/2026/07/02/mark-zuckerberg-tells-staff-that-ai-agents-havent-progressed-as-quickly-as-hed-hoped/ Hardware Hut: • AMD Ryzen AI Halo first look — PCMag: https://www.pcmag.com/news/amd-ryzen-ai-halo-first-look-giant-local-ai-power-in-a-pint-sized-box Technique Corner: • Know Your Unknowns — Thariq: https://thariqs.github.io/html-effectiveness/unknowns/ • Thariq on X: https://x.com/trq212/status/2073100352921215386 Post-Processing: • No One Escapes the Permanent Underclass — Borretti: https://borretti.me/article/no-one-escapes-the-permanent-underclass • AI Saves ~3% of Your Hours — Okane: https://okaneland.com/study/ai-productivity-roi-at-work/ Deep Dive: • State of the AI Economy — Exponential View: https://intelligence.exponentialview.co/ • Does Code Cleanliness Affect Coding Agents? (SonarSource) — arXiv: https://arxiv.org/pdf/2605.20049 Two Minutes to Midnight: • AI boom risks a financial crash — Telegraph: https://www.telegraph.co.uk/business/2026/06/28/ai-boom-risks-global-financial-crash-central-bankers-warn/ • Big Tech flips on the AI jobs wipeout — MSN: https://www.msn.com/en-us/money/careersandeducation/big-tech-has-suddenly-flipped-on-the-ai-jobs-wipeout-scenario/ar-AA27hbnR 🎙 About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. New episodes Fridays. • https://www.adipod.ai • humans@adipod.ai If something here gave you something to try on Monday, hit subscribe and drop a comment.

  7. Jul 3

    GLM 5.2 Undercuts Opus, Self-Rewriting Harness, AI Out-Persuades Humans & Prompt Injection as Role Confusion

    AI now out-argues expert human debaters, even coaching doesn't save them. Cap its word count though, and the entire edge drops to zero.This week: GLM 5.2 undercuts Opus, Xiaomi's self-rewriting harness, OpenAI's "jalapeno" chip, prompt injection as role confusion, and the clock ticking to 4:45. Co-hosts: Shimin Zhang, Dan Lasky, Rahul Yadav. ▸ GLM 5.2 — Z.ai's open-weight 750B MoE: Sonnet-to-Opus quality at ~1/3 the cost (~$5 vs $20 a build). Semgrep even had it beating raw Claude Code on security. ▸ Engineering jobs — SignalFire: software engineering was 2025's most resilient role (~55% of hires). Ramp: AI adopters grew headcount 10.2%, entry-level share slid 50%→34%. ▸ Tool Shed — Xiaomi's Harness X uses an AEGIS judge to rewrite its own scaffolding (Qwen3 5.9B, +44% planning). Ornith 1.0 does RL on the weights and the solution together. ▸ Hardware Hut — OpenAI + Broadcom's "jalapeno" inference chip: from the wafer photo, a systolic-array ASIC, six HBM stacks, cost-per-watt beating NVIDIA, ~9 months to tape-out. ▸ Post-Processing — Hackenberg et al.: AI out-argues lay people (~8pp) and trained debaters (~4.6pp). Cap it to a human's word count and the edge hits 0.0pp. It tripled charity donations. ▸ Listener Mail — Bloomberg on Silicon Valley engineers running a dozen agents at their kids' games. Dan's version is cognitive debt; ChainGuard wants managers at the 50th percentile of usage. ▸ Deep Dive — Yu, Cui & Hadfield-Menell: jailbreaks are a model mistaking your words for its own thoughts. The "wearing green" trick breaks GPT-5-mini and o4-mini. ▸ Two Minutes to Midnight — Masa Son doubts Musk's data-centers-in-space (~7% of the cost is electricity). OpenAI may delay its IPO toward $760B; Epoch AI sees capex outrun cash flow by Q3 2026. Clock 5:00 → 4:45. Chapters 00:00 Cold Open & Welcome02:29 News: GLM 5.2 Undercuts Opus at a Third the Cost10:37 News: Engineering Jobs, the Most Resilient?16:55 Tool Shed: Xiaomi's Harness X & Ornith22:35 Hardware Hut: OpenAI x Broadcom "Jalapeno" Chip27:55 Post-Processing: AI Out-Persuades Expert Humans39:09 Listener Mail: AI Anxiety in Silicon Valley46:43 Deep Dive: Prompt Injection as Role Confusion1:02:43 Dan's Rant: Token-Maxing Is Dead1:07:52 Two Minutes to Midnight: Space Data Centers, OpenAI's IPO, Capex Articles we discussed The Treadmill / News:• GLM 5.2 vs Opus — techstackups: https://techstackups.com/comparisons/glm-5.2-vs-opus/• GLM 5.2 beats Claude on cyber benchmarks — Semgrep: https://semgrep.dev/blog/2026/we-have-mythos-at-home-glm-52-beats-claude-in-our-cyber-benchmarks/• Engineering jobs are the most resilient — TechCrunch: https://techcrunch.com/2026/06/24/ai-was-supposed-to-kill-engineering-jobs-but-new-data-suggests-theyre-the-most-resilient/• Companies hire more after AI adoption — Ramp: https://ramp.com/data/heavy-ai-adopters-hire-more Tool Shed:• Xiaomi HarnessX rewrites its own scaffolding — VentureBeat: https://venturebeat.com/orchestration/xiaomis-harnessx-rewrites-its-own-ai-scaffolding-mid-task-and-smaller-models-gain-the-most• Ornith 1.0 — Deep Reinforce: https://deep-reinforce.com/ornith_1_0.html• Simon Willison on Ornith: https://simonwillison.net/2026/Jun/29/ornith/ Hardware Hut:• OpenAI x Broadcom "Jalapeno" inference chip — OpenAI: https://openai.com/index/openai-broadcom-jalapeno-inference-chip/ Post-Processing:• AI systems out-persuade expert humans (Hackenberg et al.) — arXiv: https://arxiv.org/pdf/2606.16475 Listener Mail:• AI anxiety is fueling burnout across Silicon Valley — Bloomberg: https://www.bloomberg.com/news/articles/2026-06-26/ai-anxiety-is-fueling-burnout-across-silicon-valley-s-tech-workers Deep Dive:• Prompt Injection as Role Confusion (Yu, Cui, Hadfield-Menell): https://role-confusion.github.io/ Two Minutes to Midnight:• Betting against Musk's AI vision — MSN: https://www.msn.com/en-us/money/other/why-one-of-tech-s-biggest-gamblers-is-betting-against-elon-musk-s-ai-vision/ar-AA26Fe1f• Archive mirror: https://archive.ph/UdzT4• Hyperscaler capex vs cash flow — Epoch AI: https://epoch.ai/data-insights/hyperscaler-capex-vs-cash-flow About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. New episodes Fridays. • https://www.adipod.ai• humans@adipod.ai If something here gave you something to try on Monday, hit subscribe and drop a comment. (00:00) - Cold Open & Welcome (02:29) - News: GLM 5.2 Undercuts Opus at a Third the Cost (10:37) - News: Engineering Jobs, the Most Resilient? (16:55) - Tool Shed: Xiaomi's Harness X & Ornith (22:35) - Hardware Hut: OpenAI x Broadcom "Jalapeno" Chip (27:55) - Post-Processing: AI Out-Persuades Expert Humans (39:09) - Listener Mail: AI Anxiety in Silicon Valley (46:43) - Deep Dive: Prompt Injection as Role Confusion (01:02:43) - Dan's Rant: Token-Maxing Is Dead (01:07:52) - Two Minutes to Midnight: Space Data Centers, OpenAI's IPO, Capex

  8. Jun 26

    Grok Buys Cursor, MidJourney Goes Hardware, Hermes Agent & Evaluation-Driven Development

    MidJourney — the AI image company — just quit image generation to build 50,000 spas that scan your body slice by slice. Then the week got weirder. Co-hosts: Shimin Zhang, Dan Lasky, Rahul Yadav. ▸ SpaceX buys Cursor — Elon's SpaceX (xAI/"Grok Cursor") is acquiring Cursor for $60B in Class A common stock — a ~60x multiple on ~$1B revenue, largely to buy an enterprise foothold. (Shimin: the first real sign of an AI-tool consolidation phase.) ▸ MidJourney goes hardware — the image-gen pioneer is licensing micro-ultrasound chips to build 50,000 body-scan spas (first one: SF, 2027), aiming for a billion scans a month. Fully private, no VC backers, a self-described "community research lab." Terabytes/second — ~500 hours of HD video per single second of scan. ▸ Tool Shed: Hermes Agent (Nous Research) — the plugin-maximalist opposite of a minimal harness like Pi: built-in memory, a self-learning skill loop, cron scheduling, swappable memory providers, and ~20 chat channels out of the box. Dan: "parachuting in with sixteen crates of supplies and a film crew." ▸ Is AI ruining our skills? (Nature) — physicians' precancerous-lesion detection fell from 28.4% to 22.4% once the AI tool was removed; 52 engineers scored 50% on understanding their own code with AI vs 67% without. Cognitive debt is showing up in the data. ▸ Claude Code is a video game (Provi.me) — the "one more prompt" loop that keeps you up three hours past bedtime, and why AI finally made B2B SaaS addictive. Plus the "agent dice" repo: roll a natural 20 and a stop hook makes the agent reflect and write itself a skill. ▸ Evaluation-Driven Development (Decoding AI) — treat every AI feature as a hypothesis and gate the PR on an offline eval pipeline (built on Opik) instead of unit tests. Gold-standard vs synthetic datasets, code-metric vs LLM-as-judge evaluators, and an "aggression" dial for how big a jerk your reviewer is. (Shimin: Newtonian physics → quantum mechanics.) ▸ Two Minutes to Midnight — ChatGPT slips under 50% share (46.4%; Gemini 27.7%, Claude 10.3%), Nvidia raises $25B in its first bond deal since 2021, and Ed Zitron walks OpenAI's FT-verified financials ($38.5B loss in 2025). ~2B users — one in four people on Earth; no 10x left. Clock moved up to 5:00. ⏱ Chapters 00:00 Cold Open & Welcome01:50 News: SpaceX Buys Cursor for $60B04:46 News: MidJourney Pivots to Body-Scan Spas11:45 Tool Shed: Hermes Agent (Nous Research)19:54 Post-Processing: Is AI Ruining Our Skills? (Nature)27:13 Post-Processing: Claude Code Is a Video Game35:23 Post-Processing: Evaluation-Driven Development (EDD)41:44 Two Minutes to Midnight: ChatGPT Under 50%, Nvidia Debt, OpenAI's Numbers55:06 Outro 🔗 Articles we discussed News:• SpaceX to acquire Cursor — CNBC: https://www.cnbc.com/2026/06/16/spacex-spcx-cursor-acquisition-ipo.html• MidJourney's medical pivot — MidJourney: https://www.midjourney.com/medical/blogpost Tool Shed:• Hermes Agent docs — Nous Research: https://hermes-agent.nousresearch.com/docs/ Post-Processing:• Is AI ruining our skills? Early results are in — Nature: https://www.nature.com/articles/d41586-026-01947-1• Claude Code is a video game — Provi.me: https://provi.me/cc-like-video-games• How Evaluation-Driven Development (EDD) works — Decoding AI (Paul Easton & Alejandro Aboy): https://www.decodingai.com/p/5b766861-0001-494f-a37f-4d4eb104dcfa Two Minutes to Midnight:• ChatGPT's market share slips below 50% for the first time — TechCrunch: https://techcrunch.com/2026/06/16/chatgpts-market-share-slips-below-50-for-first-time/• Nvidia seeks to raise over $25B in first bond deal since 2021 — Ars Technica: https://arstechnica.com/ai/2026/06/chipmaker-nvidia-seeks-to-raise-over-25b-in-first-bond-deal-since-2021/• Exclusive: OpenAI's financials — Where's Your Ed At (Ed Zitron): https://www.wheresyoured.at/exclusive-openai-financials/ 🎙 About ADI Pod ADI Pod (Artificial Developer Intelligence) is a weekly podcast about AI and software development for working developers. We go through hundreds of links and dozens of newsletters each week so you don't have to. Hosts: Shimin Zhang, Dan Lasky, Rahul Yadav. New episodes Tuesdays. • https://www.adipod.ai• humans@adipod.ai If something here changed your mind or gave you something to try on Monday, hit subscribe and leave a comment with what you tried.

Ratings & Reviews

5
out of 5
10 Ratings

About

Three engineer friends argue about AI so you don't have to. Shimin Zhang, Dan Lasky, and Rahul Yadav are working developers who've been watching AI transform their profession in real time, and they got opinions on the robot takeover. Every week the three get together to riff on the latest AI news, geek out over research papers, roast each other's tool choices, and occasionally have an existential crisis about whether the craft is dying or just getting weird. What you're signing up for: - AI news without the LinkedIn cringe: model drops, acquisitions, open-source drama, and the other stuff that actually matters if you write code for a living. - Technique corner: real tips from the trenches: spec-driven development, multi-agent orchestration, Claude.md tricks, and all the ways they've wasted hours so you don't have to. - Two Minutes to Midnight: the show's running AI bubble tracker, complete with circular funding diagrams, hyperscaler CAPEX math, and a doomsday clock they keep arguing about moving. - Deep dives that (occasionally) go deep: hallucination neurons, agentic memory, workflow automation economics, LLM architectures the papers nobody else is covering because they're hard. - Dan's Rant: Dan frequently gets mad about things. It's a whole thing. - The feelings segment: Yes, Shimin reads Tennyson on a tech podcast. Yes, Rahul wrote an AI-generated country song. No, they're not sorry. Three friends with strong opinions, questionable metaphors, and genuine love for the craft they're also mourning for. If you want to understand AI deeply, use it without embarrassing yourself, and laugh at the absurdity of it all, pull up a chair.

You Might Also Like