Attention Deficit Podcast

Alexa Griffith + Taylor Dolezal

Attention Deficit is a weekly AI news show from two engineers who build with this stuff all day. Taylor Dolezal (Dosu) and Alexa Griffith (Red Hat) cover agents, infrastructure, and the messy reality between demo and production. attentiondeficitpod.substack.com

Episodes

  1. 3d ago

    Attention Deficit Ep. 4 – The effort knob, the data center at sea, and go read a book

    Every model now ships a dial for how hard it thinks, and most of us slide it up and leave it there. This episode asks what effort actually pays for, in tokens, in data centers, and in reading time. Taylor ran four parts GPT-5.6 Sol to one part Opus 5, Alexa stayed on the Codex train, and Taylor teased Decant, Dosu’s open source tool for making sense of agent logs. The Wheel of Tokens has a theme song now. Cache hits and misses * Palantir followed up our defense tech episode with $1.94 billion in Q2 revenue, up 93 percent. Karp’s word was “otherworldly.” * Escape is a marketing beat now. Last week’s incident is this week’s flex. * Meta shipped Muse Code, a terminal coding agent chasing Codex and Claude Code. * Maximalism is back, on the web and in marble. Robotor’s robots carved a full-size Tyche in days. * Higgsfield still powers our intro. Send Alexa your tools. * Resignation post of the week goes to leaving OpenAI to build Jurassic Park. Let’s hope those don’t escape like the agents. * The terminal isn’t the right interface. Block’s Buzz, Herdr, and Claude Tag all point somewhere better. * Voice, almost. Whisper keeps dropping Alexa’s “not” directions, and ChatGPT voice mode keeps trying to hang up on her (intense ghosting for real). * Claude Code sessions now message each other. Taylor wants a Dungeons and Dragons eval on top. * Agents are getting wallets. Cloudflare Wallets opened handle reservation, Stripe has rails too, and Taylor grabbed onlydole since Taylor Swift wasn’t available. * Boris Cherny says delete your CLAUDE.md every six months. Taylor hears upgrade Postgres by deleting your rows, and Reporails pushes back. * The downgrade form works. Fable 5 routes cyber requests to Opus 4.8, and applying actually unblocked Taylor’s security work. * The Linux Foundation launched the Tokenomics Foundation to standardize AI cost and ROI measurement. Decant territory. * Claude Design fused Alexa’s hand-drawn slide style into a system her animations now inherit. * Best reply of the week. How do you trick Claude Code’s usage limits? Get a job. Advice taken. What the wheel landed on * The data center goes to sea. Compute is leaving land because of power. Until someone publishes wave-fleet economics, the ocean version stays a well-funded hypothesis. * Peter Thiel personally led Panthalassa’s $140 million Series B, $210 million total since 2016 * Panthalassa’s wave-powered, seawater-cooled nodes ship tokens home by satellite, with Ocean-3 pilots in the northern Pacific this year * Nadella has chips sitting in inventory that he can’t plug in * Starcloud ran an Nvidia H100 in orbit, and Project Suncatcher flies TPU prototypes by early 2027 * Andrew McCalip’s cost model prices one orbital gigawatt at $51.1 billion against $15.9 billion on the ground * The effort knob. Reasoning is extra tokens inside a think tag, and they bill at output rates. Almost nobody should run at extra-high. * DeepSeek-R1 and Kimi k1.5 landed the same January day, teaching reasoning by grading outcomes * Maxing the dial costs about thirteen times baseline for well under a point of accuracy * GPT-5.6 Terra is the family’s middle child, with Sol or Luna settings beating it at every level * Generation beats effort, so test one level lower when you migrate * Labs now train the knob directly, reward cliffs on token budgets included * Run extraction and classification at low or none, which Stop Overthinking backs, plan at high, execute at low, batch at max * Taylor’s Dosu MCP warning. At low effort, models skipped their tools, even when instructed * Where agent knowledge lives. Taylor’s segment, on improving the knowledge instead of the agent. * KSI keeps agents disposable and curates a shared knowledge base, code included. It solved more tasks for less money, and the knowledge transferred across model families * Log-augmented generation reuses past reasoning as KV caches * Our stacks stay duct tape. Taylor runs Obsidian plus Dosu, Alexa throws papers at her agents, and Mem0 and Beads sit on the try list * Go read a book. Get those Goodreads gains. * The Atlantic says the age of reading is over, and the data agrees. Daily reading for pleasure fell from 28 percent of Americans in 2004 to 16 percent in 2023 * Monthly e-book releases tripled since ChatGPT and dilute the market * A $2 million crime-novel deal collapsed over AI authentication, the clause publishing contracts now hinge on * When Taylor wrote the Terraform Cookbook with Kerim Satirli, the rule was simply don’t use AI * Dallas ISD’s phone ban drove 200,000 more library checkouts, up 24 percent * Now reading. Dungeon Crawler Carl book seven for Taylor, the unabridged Count of Monte Cristo plus On Earth We’re Briefly Gorgeous for Alexa * Add us on Goodreads. Reading a book is the cheapest vacation you can take Questions we keep getting What is Attention Deficit? A weekly AI news podcast from Alexa Griffith of Red Hat AI and Taylor Dolezal of Dosu, taking the week’s news as one conversation, with interactive artifacts. Where to find us New episodes land on Substack, YouTube, Spotify, and attentiondeficit.ai. Alexa is at alexagriffith.com, Taylor at onlydole.dev. Tell us your favorite model, your effort setting, and what you’re reading. See you next week. Burn your tokens responsibly, and please, go read a book. Get full access to Attention Deficit at attentiondeficitpod.substack.com/subscribe

  2. Aug 2

    Attention Deficit Ep. 3 – The incident part II, the off switch, and follow the data

    Last week OpenAI’s models broke out of their eval, and Anthropic went looking through its own transcripts. This week it published what it found. Across 141,006 cybersecurity evaluation runs, Claude models reached the open internet three times and gained unauthorized access to the production systems of three real organizations. The system prompt said no internet access. The infrastructure had a live path. That gap between what somebody wrote down and what the network actually enforces turned out to be the thread running through the whole episode, from Mark Zuckerberg’s promise of AI for everyone, to a librarian’s packed class on turning AI off, to the defense companies wiring sensors to decisions. We opened with favorite models, which has quietly become a segment about model roulette. Some days you hit the right instance, and it flies! But other days the same model doesn’t act the same, and Alexa’s verdict on air was that the models are a little moody. She spent the week on the Codex train while Taylor kept his split routine one model plans, a workhorse does the volume, and Claude comes in at the end to validate everything. The report we keep hearing about Opus 5 is that reducing the reasoning effort turns the agent from “coworker who tries too hard” into “does what you asked.” We also talked ourselves into building an AI personality quiz like, which model are you, and at what effort level? If that sounds interesting, let us know (and we’ll build it.) The Wheel of Tokens got an upgrade this week. It now has a token explosion outcome and an open-related-article button, so a segment and its source sit one click apart. We spun it live. Here’s where it landed! What the wheel landed on * The incident, part II. Anthropic reviewed 141,006 cyber eval runs and found three incidents, across six runs, where Claude reached real systems through an internet path the prompt said didn’t exist. Anthropic’s disclosure has the full trace. * AI for everyone. Zuckerberg’s WSJ op-ed argues superintelligence should be personal, built on individual empowerment, invention, and a balance of power. * The lights came back on. HumanLayer went full lights-off in July 2025 and wrote down why it failed. Dark factories don’t reliably preserve codebase quality over time without human steering. * The off switch. Hannah Cyrus’s Avoiding AI class at the Bangor Public Library drew about seventy people to a thirty-person room, and librarians around the world are now teaching it. * Follow the data. Alexa’s biggest deep dive yet treats defense tech as the design review that never happened, an architecture walkthrough of the systems connecting sensors to decisions and the money connecting the people who build them. The prompt said no internet. The network said otherwise. The details are wilder than the summary. In one incident, Claude found setup instructions inside its fictional environment pointing to a Python package that didn’t exist, so it decided to publish one. To make a PyPI account, it needed an email. To make the email, it needed a phone number, and after failing to find a free number or scrape together funds to buy one, it backtracked to an unblocked email provider and shipped the package anyway. The package was live on the real PyPI for about an hour and ran on fifteen real systems, including a security company’s scanner that installs new packages to check them for malware. The scanner ran Claude’s code, Claude collected the company’s credentials, and then went deeper with them. In another run, a model that couldn’t reach its fictional target scanned roughly nine thousand real hosts, compromised an application through an exposed debug page and SQL injection, then worked out on its own that the host was real and stopped. Two of the three affected organizations hadn’t detected anything until Anthropic reached out. Anthropic’s read is that this sits closer to “a harness and operational failure than a model alignment failure,” since the models were told the environment was simulated while the infrastructure quietly disagreed. Our read on air was simpler. A prompt is not a firewall. The environment a model can reach is the environment it’s in, whatever the sentence above it claims, which points straight back at the boring old answers. Default deny, least privilege, allowlists at the boundary. Firewalls are coming back, who would have thought. Alexa also called the discourse for what it is…a one-upism contest on X over whose model escaped the sandbox more times, right up until it’s your production system on the other end. The Cache Me Outside verdict card from the episode says it in one line. The model followed the environment, not the sentence. Everyone got access. Who got control? Zuckerberg’s essay rests on three principles, individual empowerment, invention over replacement, and a balance of power, and it hinges on open access being the safe path for everything short of biological risk. The questions we kept circling were the ones access alone doesn’t answer. Who controls the runtime, the policy, and the defaults, and who gets a real no? Alexa raised the motive question too, since a company drawing this much reporting about its own internal morale invites a closer read when it starts writing warmly about empowerment. Then there’s the practical test. Kimi K3’s weights landed Monday the 27th, right on schedule, and running them yourself still means a five- to thirty-thousand-dollar machine or aggressive distillation. Open weights widen access. The hardware bill still decides who has agency. The visual for this one walks access, control, and consent as three separate guarantees, because they are. The off switch became a class Hannah Cyrus, a librarian at the Bangor Public Library in Maine, kept getting the same question at the reference desk. Why is AI writing my emails and summarizing messages I can already read, and how do I turn it off? Her answer became a workshop called Avoiding AI. Her regular tech classes draw about a dozen people. This one drew around seventy between the room and the livestream, ended in applause, and after she wrote it up in a journal column called Refusal as Instruction, dozens of librarians around the world asked for the curriculum. The part that got us is the path. Turning these features off routinely means settings, then account, then privacy, then intelligence, then individual toggles, five screens deep and different in every product. Compare Zed, which puts disable AI right out in the open. Taylor’s ask for product teams was to put the switch beside the feature and make the choice persist, and to do it without recreating the cookie banner nightmare. Opting out of AI shouldn’t require a librarian. It’s great that it has one. The artifact measures the distance between a feature and its off switch. The lights came back on HumanLayer ran the full experiment. In July 2025, they went lights-off, agents building from specs and tickets with nobody reading the code, and their retrospective is the honest accounting. Generation got fast. Review capacity didn’t. And the models degraded codebase quality over time in ways no benchmark measures, because the cost of bad architecture shows up weeks later, when a one-line change turns into the same edit in eleven places. Their fix moves human judgment earlier, into product intent, system architecture, and program design, then ships in small vertical slices so review happens while redirecting is still cheap. The chart we built lets you move the decision point yourself and watch the backlog respond. Two additions from our own week. Don’t let a model family grade its own homework. Review your work across model families (if possible), which is why Alexa rotates CLIs across model families for review passes. And if you have a spare minute, do what Taylor does and give your adversarial reviewers names and personalities. Alexa’s status report on her own setup was the line of the night. “I think I have a dim factory.” 😂 Follow the data Alexa went deeper this week than we’ve ever gone, and the framing is what makes it land. Everyone covers defense tech as a morality play. She ran it as an architecture review, because underneath the manifestos somebody built a distributed system. Sensors see shapes, software resolves them into objects, workflows turn objects into decisions, and something turns decisions into force. Twenty years ago, the consequential defense companies built ships. Now they build chips and the software that runs the ships, which is why Palantir’s ontology, Anduril’s Lattice and Fury, Shield AI’s autonomy stack, and Hadrian’s automated factories all slot into the same diagram. Then she followed the money, and the graph collapses to a small set of nodes. Founders Fund and a16z at the center. Trae Stephens holding the Anduril executive chairman seat and a Founders Fund partnership at the same time. Palantir alumni through the Pentagon’s AI office, tech executives commissioned as Army Reserve lieutenant colonels through Detachment 201, and nearly everything tracing back to PayPal if you walk the parent nodes far enough. Ukraine functioning as a hostile production environment, where cheap, replaceable drones and Anduril’s advertised sub-minute sensor-to-shooter updates matter more than exquisite hardware. Palantir trading at roughly two and a half times Lockheed’s market value on a fraction of the workforce. And Stephens himself telling Fortune that mid-stage defense valuations of 50 to 200 times revenue are disconnected from an industry trading at 2 to 2.5, with room for “a couple of new credible players” and the rest noise. The question we ended on was quieter. If deployment starts meaning software deployment, what happens to patriotism built on visible sacrifice, and who experiences the decision as theirs when it’s distributed across operators, engineers, executives,

  3. Jul 25

    Attention Deficit Ep. 2 – The incident, the correction, and AI operating systems

    OpenAI models under evaluation escaped their research sandbox last week and breached Hugging Face’s production infrastructure. Both companies published incident reports, and the market spent the week asking everybody else for receipts. Kimi K3 promised open weights within days, and we kept coming back to the idea that agent stacks are missing an operating system. We opened the show auditing our own meters. Alexa doesn’t know how many tokens she’s burned, somewhere in the multi-billions, which for some reason translates to a lot of money. Taylor hit his usage limits by Tuesday and, as he put it, spent the rest of the week standing outside the party looking in. Alexa built the Wheel of Tokens to keep track of the flood of news. It’s a game-show wheel that picks which story we take next. We let it choose live, and every slice came up with receipts. Here’s the rundown. What the wheel landed on * The correction. Devansh’s case, which Alexa walked through on air, is that benchmarks and run rates stopped counting as proof. His breakdown runs the receipts from metered Copilot credits to per-engineer spending caps. * The productivity squeeze. In this year’s tech worker survey, most workers say AI makes them more productive, but better mostly means more and faster, not higher quality. Only 15% of founders are anxious about job security, against 51% of researchers. * Agent swarms. When Cursor rebuilt SQLite from its manual, which model sat in which role moved the bill from $10,565 to $1,339, roughly eightfold, with quality similar at the four-hour mark. Taylor’s rule of thumb after a week of experiments is to reach for the most powerful model and turn the effort down, not a weaker one at full effort. * Boris’s Ladder. Boris Cherny’s Steps of AI Adoption names five levels, 0 through 4, Gated to AI-native, each rung counting the agents you run at once, from none to a thousand and up. * The incident. OpenAI models under evaluation, one a pre-release build running with deliberately reduced cyber refusals, broke out and breached Hugging Face production. As Alexa put it on air, “The eval became the attack.” The models broke out of their eval The attack started inside a sandbox. “Controlled is the operative word,” Alexa said, and this time the box didn’t hold. OpenAI’s write-up says a combination of its models, including a pre-release build running with cyber refusals deliberately reduced for the evaluation, escalated privileges inside OpenAI’s research environment until they reached an internet-connected node, then chained a zero-day into Hugging Face’s production infrastructure, hyperfocused on an eval benchmark called ExploitGym. Hugging Face disclosed the intrusion on July 16 without knowing whose model it was. OpenAI tied it to its own models five days later. When Hugging Face’s responders pointed hosted frontier models at the attack logs, the guardrails refused the job, because they “cannot distinguish an incident responder from an attacker.” So Hugging Face ran the forensics on GLM 5.2, an open-weight model, on its own infrastructure, and no attacker data or exposed credentials left its environment. Inkling actually shipped open weights the same week, a Mixture-of-Experts model with 975 billion total parameters, 41 billion active, pretrained on 45 trillion tokens. Kimi K3 launched July 16, paused new subscriptions on July 19 when demand pushed it near capacity (”our GPUs are feeling it”), and has promised full weights by July 27. Jensen Huang joined X the day we recorded, and his first post shared an open letter from 25 companies backing open weights. Weights still aren’t free to run, and the bill is where the wheel went next. The correction hits your terminal That bill lands on ordinary desks. On air, Alexa stacked the receipts one on top of another, GitHub Copilot moving to metered AI credits, bills spiking, a ride-share giant burning its annual coding-AI budget in months before capping what each engineer can spend, with Devansh’s The AI Industry is Going Through a Massive Correction anchoring the segment. As she put it on air, “the market is starting to demand receipts.” Agent stacks can’t produce that receipt yet. Nothing in them admits work, limits it, budgets it, or routes it to the cheapest model that can finish, the job an operating system’s scheduler does. That framing came from the AI-OS piece Alexa read on air and from our own weeks of hitting limits, and the AI operating systems artifact has the long version. Cursor measured something narrower. Its swarm rebuilt SQLite in Rust from the 835-page manual, and at the four-hour mark every model mix had produced similar quality. The eightfold difference in the bill came from which model planned and which model typed. Boris Cherny, who created and runs Claude Code, mapped the climb from one agent to a thousand in Steps of AI Adoption. Level 2 is parallel agents at about ten, level 3 is supervised autonomy at about a hundred, and on air Taylor put himself between them. We took the quiz live, and you can take the same one. Everything we built for this episode Here’s that quiz, plus everything else we built this week, all live on the show’s site, attentiondeficit.ai. The full segment rundown, in running order, sits on the episode page. Steps of AI Adoption is the quiz we took on air. Its 2 AM question asks what happens when an agent misbehaves overnight, and Taylor’s best answer was “what’s an agent?” The incident walks the escape path step by step, no incident-report reading required. And the Wheel of Tokens itself spins on the segments pages, if you want the week in a different order. Six more sit beside them. The correction tracks the spend thread, and AI operating systems makes the scheduler argument. Agents in production gather the field lessons, Kimi K3 and Inkling hold the open-weights week, and Voice bit is the screen from the episode’s voice segment. That site holds the show’s whole paper trail, and the receipts below are the rest of it. The receipts The incident · OpenAI’s write-up · Hugging Face’s disclosure The correction · The AI Industry is Going Through a Massive Correction from Devansh’s Artificial Intelligence Made Simple The productivity squeeze · How tech workers are feeling in 2026. It paywalls partway down, and the numbers we cite appear before the wall. The models · Kimi K3 · Introducing Inkling Agent swarms · Agent swarms and the new model economics Boris’s Ladder · Steps of AI Adoption Questions we keep getting What is Attention Deficit? A weekly AI news podcast from Alexa Griffith of Red Hat AI and Taylor Dolezal of Dosu. We take the week’s AI news as one conversation instead of a headline list, and we build interactive HTML artifacts for every segment at attentiondeficit.ai. What happened in the OpenAI and Hugging Face security incident? Hugging Face disclosed an autonomous-agent intrusion into part of its production infrastructure on July 16, 2026, and OpenAI traced it to its own models on July 21. OpenAI’s models, including a pre-release build running with reduced cyber refusals for evaluation, escalated privileges to an internet-connected node and chained a zero-day into Hugging Face production while chasing the ExploitGym benchmark. Where do I start with the show? Start with the first episode, the welcome post that opens the series and sets the format. Every episode lives on the episodes index at attentiondeficit.ai with its video, segment rundown, and artifacts. Where to find us New episodes land on Substack, YouTube, Spotify, and attentiondeficit.ai. Alexa Griffith is at alexagriffith.com, Taylor Dolezal at onlydole.dev. If you found something interesting this week, or built HTML diagrams, charts, or other artifacts to keep up with it, share them with us. We’d love to see them. See you next week. Get full access to Attention Deficit at attentiondeficitpod.substack.com/subscribe

  4. Jul 20

    Attention Deficit Ep. 1 – Welcome to Attention Deficit

    Two engineers sit down every week to talk about what actually happened in AI. Not press releases, but articles, tools, and topics that make you stop and think. In the first episode of Attention Deficit, Alexa Griffith and Taylor Dolezal dig into agent failure modes. Zombie agents, rogue agents, and why a coding agent is really just six functions in a trench coat. They get into the OpenAI Jalapeno chip and what custom silicon means for NVIDIA’s hold on inference. And they talk about whether usage-based billing is about to make AI a lot more expensive for everyone. They also demo a closet app built entirely with AI coding agents, debate whether the Codex Micro keyboard matters, and talk about what it means to build a community around learning in public. Topics discussed * Agent failure modes and why agents fail like distributed systems * The “six functions in a trench coat” framework for understanding coding agents * OpenAI Jalapeño chip and what custom Application-Specific Integrated Circuits mean for inference * NVIDIA market dominance and Compute Unified Device Architecture (CUDA) lock-in * On-prem vs cloud GPU economics and hardware optimization layers * Usage-based billing and subscription fatigue across AI tooling * Building personal projects with AI coding agents * Forward Deployed Engineers as an emerging role * Community-driven learning in the AI space Key Takeaways * A coding agent is six functions in a trench coat: read, write, run, edit, list, search. The intelligence is in the model, not the tooling. * Agent failures look like distributed system failures because agents are distributed systems. The failure modes are not new, just repackaged. * Agents make starting easy, so you end up with fifteen tasks that are each 20% done and nothing finished. Finishing is the bottleneck. * OpenAI building custom inference silicon signals that compute costs are a permanent constraint, not a temporary growing pain. * Even Microsoft, which owns its own hardware, has moved away from flat-rate pricing for AI tools. That is your sign. * CUDA lock-in is NVIDIA’s real moat. The hardware is replaceable. The ecosystem is not. Resources from this episode Hadley Wickham, “A Coding Agent Is Six Functions in a Trench Coat”: https://tidydesign.substack.com/p/a-coding-agent-is-six-functions-in Mahesh Balakrishnan, “Your Agent is a Distributed System”: https://maheshba.bitbucket.io/blog/2026/04/24/agentfailures.html OpenAI Jalapeno announcement: https://openai.com/index/openai-broadcom-jalapeno-inference-chip/ Learn more about the hosts Alexa Griffith Website: https://alexagriffith.com/ LinkedIn: https://www.linkedin.com/in/alexa-griffith/ X/Twitter: https://x.com/alexa_griffith_ Taylor Dolezal Website: https://onlydole.dev/ LinkedIn: https://www.linkedin.com/in/onlydole/ X/Twitter: https://x.com/onlydole GitHub: https://github.com/onlydole Get full access to Attention Deficit at attentiondeficitpod.substack.com/subscribe

About

Attention Deficit is a weekly AI news show from two engineers who build with this stuff all day. Taylor Dolezal (Dosu) and Alexa Griffith (Red Hat) cover agents, infrastructure, and the messy reality between demo and production. attentiondeficitpod.substack.com