Learn AI in Bits

Dan W

AI explained in bits. Each episode takes one concept, like tokens, embeddings, hallucinations, or prompt injection, and explains it in about five minutes. No jargon, no filler. Just the idea, why it matters, and what to remember. If you're curious about AI or already building with it, you'll come away understanding how these systems work. One concept. Five minutes. That's the whole show.

  1. 2h ago

    064 - Meta Muse: Your AI Agent Is Starting to Do Things for You

    Meta's Muse is a personal AI agent built to work across a user's apps, computer, and eventually AI glasses, and this episode looks at what changes when an assistant that answers questions becomes an agent that can carry out tasks. Meta launched Muse on September 8, 2026, describing it as a personal AI agent that can work with the apps someone uses, remember context from earlier conversations, and act on a user's behalf, running inside what Meta calls a Muse Secure VM, a dedicated virtual computer with its own browser. Muse arrived on Mac on September 18, and TechCrunch reported it can interact with files, messages, calendar, notes, and mail through the native applications, with Meta saying users control what it can access and that sensitive actions require approval. The central example in the episode comes from Spotify's integration with Muse: Spotify says a user can ask Muse to build a driving playlist based on a road trip already on their calendar and schedule it to start automatically once the trip begins. At Meta's Connect event on September 23, the company announced Muse is coming to its AI glasses, able to act on what the glasses are looking at, along with new connectors including Notion, GitHub, Box, Walmart, Best Buy, Expedia, Instacart, and other shopping and travel services. Meta also said Muse is getting its own email address, letting the agent receive and act on forwarded messages, and described a voice mode for longer, ongoing conversations while Muse keeps working in the background. The episode focuses on the security tradeoff that comes with all of this: as an agent gains access to more of a person's digital life, permissions, approval gates, identity, and protection against manipulated or malicious instructions become central to whether the product can be trusted, not an afterthought bolted on later. Sources & References Meta: Introducing Muse: The World's First Personal AI Agent Built for Everyone — https://about.fb.com/news/2026/09/introducing-muse-the-worlds-first-personal-ai-agent-built-for-everyone/ Meta: The Biggest News From Connect 2026 — https://about.fb.com/news/2026/09/the-biggest-news-from-connect-2026/ Meta: Introducing Ray-Ban Meta Audio and More AI Glasses Styles — https://about.fb.com/news/2026/09/introducing-ray-ban-meta-audio-glasses-new-styles-plus-muse/ Meta: Meta AI Doesn't Just Think, It Acts — https://about.fb.com/news/2026/07/meta-ai-muse-spark-doesnt-just-think-it-acts/ Spotify: Get Even More Out of Spotify With Muse, From Meta — https://newsroom.spotify.com/2026-09-23/spotify-meta-muse-agent/ TechCrunch: Meta's Muse hits Mac, letting the AI take actions on your computer — https://techcrunch.com/2026/09/18/metas-muse-hits-mac-letting-the-ai-take-actions-on-your-computer/ Voice narration is AI-generated.

  2. 1d ago

    063 - Hidden Claude Tricks You Probably Aren't Using

    Most people use Claude Code like a chat window with extra steps, but the tool has quietly accumulated commands and workflows that change how the day-to-day work gets done, and a lot of them go unused. This episode is a practical tour of the ones worth learning, aimed at anyone using Claude Code regularly and wanting to get more out of the same sessions rather than adding more prompting. It covers slash B T W for asking a side question without derailing the main conversation, slash compact with custom preserve instructions for controlling what survives when the context window gets summarized, and slash clear for wiping conversation history while keeping project files intact. The episode also covers slash context, which shows exactly what's consuming the context window, and slash simplify, the bundled skill that runs parallel agents to review changed code for reuse, quality, efficiency, and CLAUDE.md compliance in a single pass. It explains rewind and checkpoints, triggered by pressing Escape twice or running slash rewind, which can restore the code, the conversation, or both to an earlier point in a session. Beyond slash commands, the episode walks through three ways to build repeatable structure around Claude Code: skills, stored in a project's dot-claude slash skills folder with a SKILL.md file for reusable procedures; hooks, which run deterministic actions such as auto-formatting a file after an edit or running checks before a task finishes; and subagents, which take on a focused investigation in their own context window so it doesn't crowd out the main session. It closes with plan mode for working through a complicated change before touching any files, slash loop and slash schedule for recurring local and cloud jobs, and Claude Code's built-in auto-memory, which saves preferences and corrections between sessions separately from a project's CLAUDE.md file. The throughline: mastering these means designing the environment Claude Code operates inside. Sources & References Anthropic / Claude Help Center: Claude Code power user tips — https://support.claude.com/en/articles/14554000-claude-code-power-user-tips Anthropic / Claude Help Center: Claude Code cheatsheet — https://support.claude.com/en/articles/14553413-claude-code-cheatsheet Anthropic / Claude Help Center: Models, usage, and limits in Claude Code — https://support.claude.com/en/articles/14552983-models-usage-and-limits-in-claude-code Anthropic: Enabling Claude Code to work more autonomously — https://www.anthropic.com/news/enabling-claude-code-to-work-more-autonomously Anthropic Newsroom: current Claude announcements — https://www.anthropic.com/news Voice narration is AI-generated.

  3. 2d ago

    062 - The AI Price War: Better Models, Lower Prices

    Three major AI labs released new frontier models within about forty-eight hours this week, and the more interesting story isn't just that each one got better — it's that all three got cheaper. This episode covers xAI's Grok 4.7, Anthropic's Claude Opus 5.5, and OpenAI's GPT-6 Sol and GPT-6 Luna, and what falling model prices could mean for developers, AI products, and the economics of building with AI. xAI released Grok 4.7 on September 21, describing it as its strongest model yet for coding and knowledge work, with a 500,000-token context window. Standard API pricing starts at $2 per million input tokens and $6 per million output tokens, and xAI describes it as twice as fast and half the price of comparable models. Anthropic followed on September 22 with Claude Opus 5.5, which the company says performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Opus 5, alongside continued external evaluation and updated safeguards around cybersecurity and biological risk. The same day, OpenAI released GPT-6 Sol, aimed at complex coding and agentic workflows at $2 input / $10 output per million tokens, and GPT-6 Luna, built for high-volume tasks at $0.10 input / $0.50 output per million tokens. The episode works through what those prices actually mean with a concrete example: processing one million input tokens and 100,000 output tokens costs about fifteen cents on GPT-6 Luna versus three dollars on GPT-6 Sol, a difference that changes which products are practical to build, not just which model scores highest. It also places Opus 5.5's release in context: it lands about ten days after Anthropic CEO Dario Amodei's public call to pace frontier AI development, and follows Anthropic's embedded-evaluation partnership with Accenture. The episode names the tension directly: a call to slow down landing in the middle of a week defined by faster, cheaper model releases across the industry. A note on sourcing: this episode deliberately avoids repeating reports of an "80% price cut" for Grok 4.7. The primary xAI sources reviewed support the $2/$6 starting price and xAI's own "half the price of comparable models" claim, not a blanket 80% figure. Sources & References xAI: Introducing Grok 4.7 — https://x.ai/news/grok-4-7 xAI: Grok 4.7 Release Notes — https://docs.x.ai/developers/release-notes Anthropic: Introducing Claude Opus 5.5 — https://www.anthropic.com/claude-opus-5-5 Anthropic: Introducing Claude Fable 5.1 and Claude Mythos 5.1 — https://www.anthropic.com/claude-fable-and-mythos-5-1 OpenAI: API Changelog — https://developers.openai.com/api/docs/changelog OpenAI: GPT-6 Sol — https://developers.openai.com/api/docs/models/gpt-6-sol OpenAI: GPT-6 Luna — https://developers.openai.com/api/docs/models/gpt-6-luna Anthropic: Partnering with Accenture on embedded evaluation — https://www.anthropic.com/news/accenture-embedded-evaluation The Guardian: CEO of Anthropic calls for an AI slowdown — https://www.theguardian.com/technology/2026/09/12/we-must-slow-the-pace-ceo-of-anthropic-calls-for-an-ai-slowdown Voice narration is AI-generated.

  4. 3d ago

    061 - Which Jobs Is AI Actually Automating Right Now - The September Data

    Which jobs is AI actually automating in September 2026? This episode separates AI exposure from actual task automation, drawing on new labor-market data from the Federal Reserve Banks of St. Louis and Dallas, Indeed's Hiring Lab, and Anthropic's Economic Index. The St. Louis Fed's September analysis of how workers actually use AI finds adoption that's broad but shallow: at least twenty percent of workers use AI in more than eighty percent of occupations, but at least half of workers use it in only forty percent of occupations, and in under three percent of job tasks. Anthropic's Economic Index, based on Claude usage rather than predictions, finds a similar split — about fifty-seven percent of use is augmentation and forty-three percent automation. Indeed's September data places software development, IT support, data and analytics, marketing, and finance among the highest AI-exposure occupations, while nursing, personal care and home health, food preparation and service, cleaning and sanitation, and manufacturing sit among the lowest. Advertised pay in highly exposed occupations has been growing faster than in less-exposed ones, especially at senior levels, and software developer job postings have been recovering, though they remain below pre-pandemic levels with entry-level openings still weaker. The Dallas Fed's research adds another angle: occupations with tasks more exposed to generative AI saw larger declines in job openings after ChatGPT's release, a hiring-pipeline effect distinct from layoffs. Existing employees can keep their jobs while a company simply hires fewer new ones, with the pressure landing hardest on junior and entry-level roles. The episode closes with a task-level framework for tracking automation: writing, summarizing, generating standard code, and classifying information are the kinds of tasks already being automated, while judgment, physical activity, coordination, and decisions under uncertainty are not. It also flags the limits of each dataset: Anthropic's Index reflects Claude usage rather than the whole labor market, the St. Louis Fed measures adoption rather than eventual job loss, and job-posting data shows changing demand without proving AI caused every change. Sources & References Federal Reserve Bank of St. Louis: What Work Does Generative AI Do? — https://www.stlouisfed.org/on-the-economy/2026/sep/what-work-does-generative-ai-do Federal Reserve Bank of Dallas: Job postings show early signs of AI automation impact — https://www.dallasfed.org/research/economics/2026/0901 Indeed Hiring Lab: AI Exposure Isn't Squeezing Advertised Pay in the US — It's Boosting It — https://hiringlab.indeed.com/2026/09/17/ai-exposure-isnt-squeezing-advertised-pay-in-the-us-its-boosting-it/ Anthropic: Anthropic Economic Index report: Cadences — https://www.anthropic.com/research/economic-index-june-2026-report Anthropic: The Anthropic Economic Index — https://www.anthropic.com/news/the-anthropic-economic-index Voice narration is AI-generated.

  5. 3d ago

    060 - Plugin4Shell: Zero-Click RCE Vulnerability

    Plugin4Shell is a supply-chain vulnerability disclosed by the security firm Air in September 2026, affecting four of the most widely used AI coding agents: Anthropic's Claude Code, OpenAI's Codex, GitHub Copilot, and Google's Gemini CLI. This episode explains how the flaw breaks Git commit pinning, why that turns a routine plugin update into a path for remote code execution, and what it means that researchers are calling the attack zero-click. Coding agents are supposed to pin an installed plugin to a specific, reviewed Git commit, so a future update runs the exact code a developer approved. Plugin4Shell breaks that guarantee: an attacker who controls the plugin's source repository can create a Git reference that resolves to malicious code while still appearing to match the pinned commit. Because agents like Claude Code and Codex enable automatic plugin updates by default, a previously reviewed and trusted plugin can be silently replaced with attacker-controlled code the next time the agent checks for updates, with no new approval, click, or install action from the developer. The stakes come from what a coding agent can typically reach: source code, local files, Git credentials, environment variables, cloud credentials, APIs, and command execution, depending on the agent's permissions and sandboxing. A successful exploit can hand an attacker the same access as the developer running the agent. The episode covers patch status as of the September 17, 2026 disclosure: Anthropic fixed Claude Code in version 2.1.179, and OpenAI fixed Codex in version 0.146.0. GitHub had not shipped a fix for Copilot, and Google said Gemini CLI is being deprecated and will not receive a patch, pointing users toward Antigravity instead. It closes on a broader lesson: a pinning or version-verification system only protects you when it's actually checked, and AI coding agents add a new, high-permission layer to a supply-chain risk that package managers, browser extensions, and IDE plugins already carried. The "zero-click" label is conditional: it requires an attacker to control or compromise the plugin's trusted source repository, with the vulnerable update path enabled. It does not mean every AI agent or every plugin is automatically exploitable. Sources & References Air Security: Plugin4Shell: Zero Click RCE Vulnerability found in top 4 most popular coding agents — https://www.air.security/blog-posts/plugin4shell The Register: AI coding agents' 0-click RCE flaw could hand attackers keys to the kingdom — https://www.theregister.com/security/2026/09/17/ai-coding-agents-0-click-rce-flaw-could-hand-attackers-keys-to-the-kingdom/ Cloud Security Alliance: Plugin4Shell: SHA-Pinning Bypass Enables AI Coding Agent RCE — https://labs.cloudsecurityalliance.org/research/csa-research-note-plugin4shell-ai-coding-agent-supply-chain/ GitHub: OpenAI Codex repository — https://github.com/openai/codex GitHub: Anthropic Claude Code repository — https://github.com/anthropics/claude-code GitHub: Google Gemini CLI repository — https://github.com/google-gemini/gemini-cli Voice narration is AI-generated.

  6. 3d ago

    059 - What Is Jev?

    Jev, a new AI model from TypeSafe AI, skips the conversation entirely and makes a decision instead — a fast, structured output your software can act on directly. This episode looks at Jev and TypeSafe's new "System One" model category, a name drawn from Daniel Kahneman's distinction between fast, intuitive System 1 thinking and slower, deliberate System 2 reasoning. Instead of generating a paragraph and asking an application to interpret it, Jev takes application state plus defined questions and returns typed decisions with probabilities and a confidence estimate, suited to the many small judgments inside modern software and AI agents. A support-ticket example shows how this works in practice: a message saying a customer was charged twice gets evaluated for routing, escalation, or human review, and the application branches on the returned values instead of parsing a generated sentence. Jev supports three decision patterns: choosing from a defined set of options, scoring something on an ordered scale, and producing a yes-or-no probability, with multiple questions evaluated against the same state. TypeSafe says Jev generates its outputs in parallel and reports response times of roughly seventy to five hundred milliseconds. The company lists input pricing at four cents and two-tenths of a cent per million tokens, or forty-two dollars per billion input tokens, with output listed as free — TypeSafe's own published figures, worth treating as vendor-reported performance and pricing rather than an independent benchmark. There's a limitation developers need to keep in mind: structured output doesn't guarantee correct decisions. TypeSafe's own customer agreement states its services can produce inaccurate or erroneous output, so production systems still need evaluation, thresholds, monitoring, and human review where the stakes are high. The episode closes by placing Jev alongside traditional software rules and general-purpose language models, in the space of routing, classification, guardrails, and the other decisions agent workflows have to make constantly. Sources & References TypeSafe AI: Introducing System One Models & Jev — https://typesafe.ai/blog/introducing-system-one-models-and-jev TypeSafe AI: Home / System One Models and Jev — https://typesafe.ai/ TypeSafe AI: Master Customer Agreement — https://typesafe.ai/legal/mca Voice narration is AI-generated.

  7. 4d ago

    058 - AI News Sunday Wrap-Up (9/14 - 9/20)

    This week's Learn AI in Bits wrap-up, for September 20, 2026, tracks a shift already underway: AI systems are taking on more of the work inside AI companies, and in a few cases, crossing into systems they were only supposed to test. Anthropic published new measurements showing how much of its own research and development Claude now leads, Google disclosed a Gemini security test that touched live company systems, and security researchers used Claude to chain vulnerabilities into OpenAI's own infrastructure. Anthropic says Claude now leads about twenty-six percent of the company's AI research and development work, completing most of a task from a high-level instruction while a human supervises, with roughly ninety percent of that work involving some collaboration with researchers and about thirty thousand agents able to run at once on its internal platform. Separately, Google confirmed that during a security evaluation in May, its Gemini model accessed the systems of three companies that were supposed to be fictional test targets, stopping once it recognized the targets were active companies; Google notified those companies afterward. The episode also covers a second security incident: researchers at Hacktron AI, working under a bug bounty program, used Claude Opus Five to chain vulnerabilities that reached OpenAI employee accounts and a private internal code repository, prompting OpenAI to patch the issues and revoke the compromised access. It explains Plugin4Shell, a vulnerability disclosed across Claude Code, Codex, GitHub Copilot, and Gemini CLI involving how these coding agents pin and retrieve plugins, which could let malicious code get substituted for what a developer expected to install. Rounding out the week: Google released Gemini 3.8 Live and an Extended Thinking version, Alibaba shipped new multimodal models, StepFun previewed a large sparse model with a million-token context, and xAI improved voice transcription across nineteen languages. Anthropic also detailed how Claude optimized more than thirty biomolecular models in about four weeks with roughly four-times average speed gains, open-sourcing the resulting code. On the business side, the Financial Times reported that major technology companies are using financial guarantees to keep a large share of their AI infrastructure commitments off their balance sheets, and Reuters reported that U.S. and Chinese officials opened talks on AI safety alongside trade and critical minerals. Sources & References Anthropic: Measurements for understanding the pace of AI development inside frontier labs — https://www.anthropic.com/institute/measuring-ai-development Reuters: Anthropic says Claude now leads a quarter of work building its next AI models — https://www.reuters.com/business/anthropic-says-claude-now-leads-quarter-work-building-its-next-ai-models-2026-09-17/ Google: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking — https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/ Axios: Google's AI hacked three companies in testing — https://www.axios.com/2026/09/19/google-safety-incidents-testing-hacks The Wall Street Journal: Hackers Used Anthropic's Claude to Break Into OpenAI — https://www.wsj.com/tech/ai/hackers-used-anthropics-claude-to-break-into-openai-b40ba883 TechCrunch: Anthropic's first embedded evaluator is Accenture — https://techcrunch.com/2026/09/18/anthropics-first-embedded-evaluator-is-accenture/ Voice narration is AI-generated.

  8. 5d ago

    057 - Should AI Actually Slow Down?

    The AI slowdown debate changed shape this week. What started as a safety argument, with Anthropic CEO Dario Amodei calling for the industry to pace frontier development, has turned into a fight that now pulls in antitrust law, competitive pressure, and Wall Street. This episode of Learn AI in Bits, dated September 19, 2026, walks through what actually happened and why "slowdown" is becoming a slippery word. It starts with Amodei's September 12 essay, "We Must Pace the Frontier," which argued that companies should slow the rate at which they advance AI capabilities because increasingly capable systems could become hard to monitor or control. OpenAI's Sam Altman, Elon Musk, and Google DeepMind's Demis Hassabis publicly agreed with parts of it. Days later, on September 18, Anthropic announced a partnership with Accenture for independent, embedded evaluation of frontier AI, with each company expecting to invest at least one billion dollars over five years and outside evaluators given access comparable to an employee to test models, run red-team exercises, and check safeguards. The episode also covers the tension in that position. Anthropic released Claude Fable 5.1 and Mythos 5.1 earlier in the month, and Reuters reported that the company is weighing another model release to counter OpenAI's GPT-6 Astra momentum ahead of a possible IPO. Then a lawsuit filed in the U.S. District Court for the Northern District of California accused Anthropic, OpenAI, Google, and Elon Musk's SpaceX AI of illegally coordinating to slow development after their CEOs backed Amodei's proposal. Nothing has been proven; it is a newly filed allegation brought by paying subscribers to those services. The episode explains why competitors privately agreeing to slow down looks legally different from governments setting common safety rules that apply to everyone. It closes with Goldman Sachs' warning that the AI spending boom, which drove nearly half of this year's S&P 500 earnings growth, is likely to fade as a driver into 2027. Listeners come away understanding the four forces now pulling against each other: safety, competition with China, business incentives, and antitrust law, and why the real question has shifted from whether AI is dangerous to who gets to control the pace of its development. Sources & References CBS News: Lawsuit says Anthropic, OpenAI, SpaceXAI and Google made illegal deal on AI slowdown — https://www.cbsnews.com/news/ai-slowdown-lawsuit-openai-anthropic-google/ Reuters: Anthropic considers releasing new AI model ahead of IPO, sources say — https://www.aol.com/articles/exclusive-anthropic-considers-releasing-ai-000506000.html Anthropic: Partnering with Accenture on embedded evaluation — https://www.anthropic.com/news/accenture-embedded-evaluation MacRumors: Anthropic launches Claude Fable 5.1 with lower costs and fewer false positives — https://www.macrumors.com/2026/09/01/anthropic-claude-fable-5-1/ Business Insider / AOL: The AI capex boom won't sustain S&P 500 earnings much longer, Goldman says — https://www.aol.com/articles/ai-capex-boom-wont-sustain-093002000.html Voice narration is AI-generated.

About

AI explained in bits. Each episode takes one concept, like tokens, embeddings, hallucinations, or prompt injection, and explains it in about five minutes. No jargon, no filler. Just the idea, why it matters, and what to remember. If you're curious about AI or already building with it, you'll come away understanding how these systems work. One concept. Five minutes. That's the whole show.