Learn AI in Bits

Dan W

AI explained in bits. Each episode takes one concept, like tokens, embeddings, hallucinations, or prompt injection, and explains it in about five minutes. No jargon, no filler. Just the idea, why it matters, and what to remember. If you're curious about AI or already building with it, you'll come away understanding how these systems work. One concept. Five minutes. That's the whole show.

  1. 2h ago

    077 - What Is AI Training?

    How does an AI model learn in the first place? This episode builds on the show's earlier episodes on tokens, embeddings, attention, and Transformers to explain training: the process that turns a model's huge collection of adjustable parameters into something useful. For a language model, training centers on a surprisingly simple task: predicting the next token. Using the illustrative example "The cat sat on the...", the episode shows a model producing probabilities across possible next tokens, with the gap between its prediction and the correct answer measured by a loss function. Lower loss means a better prediction, and backpropagation is how the model figures out which of its millions or billions of parameters contributed to that error and in which direction each one should change, by calculating gradients. The episode explains how frameworks like PyTorch use automatic differentiation to calculate those gradients through a model's computation graph via the chain rule, instead of deriving equations by hand, and how an optimizer such as gradient descent then nudges each parameter in the direction that reduces loss, using the metaphor of moving downhill across a loss landscape. It walks through the full loop, batch by batch: predict, measure the loss, backpropagate, update parameters, repeat, with an epoch marking one complete pass through the training data. It also draws a clear line between training and inference: during training the model's errors change its parameters, while during inference those parameters are fixed, so a chatbot responding to a question is running on weights it already learned rather than learning from that individual conversation. The episode closes on two practical points: models can overfit, performing well on training data but worse on new data, and data quality shapes outcomes as much as scale does, since bad, duplicated, biased, or contaminated data can produce bad results regardless of model size. Sources & References PyTorch: The Fundamentals of Autograd — https://docs.pytorch.org/tutorials/beginner/introyt/autogradyt_tutorial.html PyTorch: Automatic Differentiation with torch.autograd — https://docs.pytorch.org/tutorials/beginner/basics/autogradqs_tutorial PyTorch: Optimizing Model Parameters — https://docs.pytorch.org/tutorials/beginner/basics/optimization_tutorial.html PyTorch: Quickstart — https://docs.pytorch.org/tutorials/beginner/basics/quickstart_tutorial.html Li et al.: Mechanics of Next Token Prediction with Self-Attention — https://arxiv.org/abs/2403.08081 Voice narration is AI-generated.

  2. 1d ago

    Quick AI News - 10/06/2026 -Reflection AI's Beam, ChatGPT Ads, and AI Becoming Infrastructure

    This quick news update covers the AI developments from October 5 into October 6, 2026 worth knowing, from a new open-weight model to a critical security patch and a shift in how big tech companies use Anthropic's Claude internally. Reflection AI released details of Beam, its first open-weight model: a mixture-of-experts design with 501 billion total parameters but only about 23 billion active per token. Reflection says Beam is competitive with models like GLM 5.2 on coding and reasoning while using roughly 3 to 4 times less inference compute, though that efficiency comparison is the company's own claim rather than an independently verified benchmark. Weights are expected to follow later in October under an Apache 2.0 license. OpenAI announced a new advertising format for ChatGPT, which it says now reaches 1.2 billion people each week, with visual ads set to begin testing in the US later this month. The ads can appear alongside image generation, won't change model responses or user-generated images, and won't show for Plus, Pro, or Enterprise subscribers. On security, GitLab patched a critical vulnerability in its self-hosted AI Gateway, CVE-2026-90970, with a CVSS score of 9.9, that could let an authenticated user with Duo Agent Platform access escape the prompt-template sandbox and execute commands on the Gateway. Separately, researchers using Anthropic's Mythos model identified a vulnerability in the open-source Rejetto HFS file server involving weak random-number generation, enabling session forgery and remote code execution, with attackers probing it shortly after disclosure. On the business side, The Information reports Microsoft and Meta are both cutting internal use of Claude: Microsoft's expected internal Claude spending reportedly fell by more than a third, and Meta's internal Claude Code users reportedly dropped from about 60,000 to around 30,000, as both companies push employees toward their own AI tools. Separately, two former Groq engineers sued over Nvidia's roughly $20 billion 2025 deal with Groq, alleging the transaction favored insiders at other shareholders' expense. Sources & References Reflection AI: Introducing Beam: Reflection's 501B open-weight model — https://reflection.ai/blog/introducing-beam OpenAI: Building advertising for the way people use AI — https://openai.com/index/new-chatgpt-ads-format-and-measurement/ GitLab: Critical Patch Release for Self-Hosted AI Gateway — https://docs.gitlab.com/releases/patches/other-patches/patch-release-gitlab-ai-gateway-19-4-1-released/ Horizon3: Anthropic's Mythos and Rejetto HFS — https://horizon3.ai/attack-research/disclosures/anthropic-mythos-rejetto-hfs-rce/ The Register: Anthropic's Mythos vulnerability story — https://www.theregister.com/security/2026/10/03/anthropics-super-bug-hunting-model-mythos-is-hardcore-good-at-math-as-latest-vuln-under-attack-shows/ The Information: Microsoft Slashes Internal Claude Spending by a Third — https://www.theinformation.com/articles/meta-microsoft-work-wean-staff-anthropics-claude Financial Times: Nvidia's $20bn Groq Deal Faces Lawsuit — https://www.ft.com/content/93ee425d-9ac7-4548-8cc9-fef2e0670787 Voice narration is AI-generated.

  3. 1d ago

    076 - What Is an AI Transformer [revised episode]?

    What is a Transformer, and why did this one architecture become the foundation for so much of modern AI? This episode builds on the show's earlier episodes on tokens, embeddings, and attention to explain how a Transformer puts those pieces together into a working system. A Transformer is a neural network architecture designed to process sequences, introduced in the 2017 paper "Attention Is All You Need" for machine translation and later adopted far beyond it. Tokens become embeddings, positional information is added so the model knows where each token sits in the sequence, and the result moves through repeated layers combining attention with a feed-forward network, residual connections, and layer normalization, so later layers build on patterns found by earlier ones. The episode uses the original 2017 base model as a concrete example: six encoder layers and six decoder layers, eight attention heads, a 512-dimensional representation per token, and a feed-forward network that expands to 2,048 dimensions before compressing back down. It walks through that model translating English to French, with the encoder building contextual representations and the decoder generating output one token at a time while a mask stops it from looking ahead at tokens it hasn't produced yet. It also draws the distinction between that original encoder-decoder design and the decoder-only Transformers behind modern chatbots like GPT, which skip the encoder entirely and use causal attention to predict the next token from what came before. The episode explains why the architecture scaled so well in the first place: training can process many token positions in parallel instead of one at a time, which fits how modern hardware handles large matrix operations, though standard self-attention's computational cost still grows roughly with the square of sequence length. It closes by stressing that a Transformer is an architecture, not a single model: GPT, BERT, and many other systems are all built from Transformer ideas, using different parts of it for different purposes. Sources & References Vaswani et al.: Attention Is All You Need — https://arxiv.org/abs/1706.03762 PyTorch: Accelerating PyTorch Transformers by replacing nn.Transformer with Nested Tensors and torch.compile() — https://docs.pytorch.org/tutorials/intermediate/transformer_building_blocks.html PyTorch Tutorial: Implementing High-Performance Transformers with Scaled Dot Product Attention — https://docs.pytorch.org/tutorials/intermediate/scaled_dot_product_attention_tutorial Hugging Face: How do Transformers work? — https://huggingface.co/docs/course/chapter1/4 Hugging Face: Transformer Architectures — https://huggingface.co/docs/course/chapter1/6 Voice narration is AI-generated.

  4. 1d ago

    075 - What is AI Attention?

    When an AI model reads a sentence, it doesn't treat every word as equally important. This episode explains attention, the mechanism that lets a model decide which parts of its context deserve more weight when computing what comes next, and one of the central ideas behind every modern Transformer-based language model. The episode walks through the basic calculation using three learned representations: query, key, and value. A query represents what a token is looking for, a key represents what another token offers for matching, and a value carries the information passed along if that token turns out to be relevant. The model compares a query against keys using a similarity score, scales those scores, and passes them through softmax so they turn into weights that sum to one. Those weights combine the value vectors into a new, context-aware representation, using the sentence "The animal didn't cross the road because it was tired" to show how surrounding words give the model clues about what a pronoun like "it" refers to. It also covers self-attention, where queries, keys, and values all come from the same sequence, and multi-head attention, where a Transformer runs several attention calculations in parallel so each head can specialize in a different kind of relationship in the data. A key constraint for text generation is causal masking, which stops a model from attending to tokens it hasn't generated yet, preserving the left-to-right nature of generation. Because full attention creates a relationship between every pair of positions in a sequence, its computational cost grows roughly with the square of sequence length, which is why longer context windows demand more compute and memory. The episode connects this to the engineering solutions that make it practical, citing PyTorch's scaled dot product attention and optimized kernels like FlashAttention, and traces the idea back to the 2017 paper "Attention Is All You Need," which introduced a Transformer architecture built entirely around attention instead of recurrence or convolution. The episode closes by tying attention to the show's earlier episodes on tokens and embeddings: tokens provide the pieces, embeddings turn them into numbers, and attention is what lets those numbers interact. Sources & References Vaswani et al.: Attention Is All You Need — https://arxiv.org/abs/1706.03762 PyTorch Documentation: scaled_dot_product_attention — https://docs.pytorch.org/docs/main/generated/torch.nn.functional.scaled_dot_product_attention.html PyTorch Tutorial: Implementing High-Performance Transformers with Scaled Dot Product Attention — https://docs.pytorch.org/tutorials/intermediate/scaled_dot_product_attention_tutorial PyTorch Tutorial: Accelerating PyTorch Transformers with Transformer building blocks — https://docs.pytorch.org/tutorials/intermediate/transformer_building_blocks.html Voice narration is AI-generated.

  5. 2d ago

    074 - [audio fixed] AI News Sunday Wrap-Up (9/28 - 10/4)

    The AI industry packed an unusual amount into a single week, and this Sunday wrap-up covers the developments from September 28 through October 4, 2026 that say the most about where the industry is heading. Anthropic opened the week on September 28 by releasing Claude Sonnet 5.5, which Anthropic says runs more than 30% faster than Sonnet 5 and costs up to 30% less for most work, with a large reported jump on Terminal-Bench, an agentic-coding evaluation. The next day, OpenAI held its DevDay, introducing GPT-6.1 Sol for coding, computer use, and professional work, a public beta of its Agents API with hosted execution and multi-agent capability, and dots, always-on agents that keep working between conversations on their own cloud computer. OpenAI also added computer use to the Agents API, letting an agent operate software through its interface rather than a dedicated API, and GitHub followed on October 1 by putting Copilot's own computer-use feature into public preview for Windows and macOS. Google released Gemini 4 Argon on September 30, its new frontier model with a 1-million-token context window, initially available through a program for trusted cyber defenders. The same week, NVIDIA announced an Open Agent Safety Platform combining OpenShell, an open-source runtime that traces and enforces policy on agent actions, with Sentry, a watchdog that quarantines agents that move outside their boundaries. On the policy side, President Trump signed an executive order on September 29 titled "Inaugurating the Era of Super Intelligence," directing federal agencies to define the term "Super Intelligence," or SI, and to replace "artificial intelligence" with it in official communications. The administration separately announced a voluntary agreement with major AI companies covering internal controls, outside auditing, and board oversight of AI safety. Taken together, the episode argues the week marks a shift from model releases toward full systems: agents that remember tasks, use tools, operate software, and keep working after a person steps away, raising new engineering problems around permissions, sandboxing, and control. Sources & References Anthropic: Claude Sonnet 5.5 — https://www.anthropic.com/claude-sonnet-5-5 OpenAI API Changelog: September 2026 updates — https://developers.openai.com/api/docs/changelog OpenAI: DevDay 2026 — https://devday.openai.com/ OpenAI Help Center: ChatGPT release notes — https://help.openai.com/en/articles/6825453-chatgpt-release-notes Google: September 2026 AI updates, including Gemini 4 Argon — https://blog.google/innovation-and-ai/technology/ai/google-ai-updates-september-2026/ GitHub: Copilot computer use public preview — https://github.blog/changelog/2026-10-01-github-copilot-can-now-interact-with-desktop-apps/ NVIDIA: Open Agent Safety Platform — https://investor.nvidia.com/news/press-release-details/2026/NVIDIA-Launches-Open-Agent-Safety-Platform-to-Secure-Agents-From-Testing-to-Deployment/ The White House: Executive Order 14434, "Inaugurating the Era of Super Intelligence" — https://www.whitehouse.gov/presidential-actions/2026/09/inaugurating-the-era-of-super-intelligence/ Voice narration is AI-generated.

  6. 3d ago

    073 - What Is an Embedding?

    What exactly is an embedding, and why is it such an important building block in modern AI? This episode explains how text becomes a numerical vector, how comparing those vectors enables semantic search, and how the same idea powers retrieval augmented generation, recommendations, and multimodal retrieval. An embedding model turns a piece of text into a vector, a list of numbers that represents patterns in the input rather than individual words or a fixed dictionary of meanings. Related sentences end up with vectors that are mathematically close together, while unrelated ones end up farther apart, in a space that can span hundreds or thousands of dimensions. As a concrete example, Hugging Face's sentence-transformers/all-MiniLM-L6-v1 model produces a 384-dimensional embedding for a single sentence, straight from its model card. The episode walks through how similarity is measured, including cosine similarity, using three example sentences to show how two related sentences land close together in the embedding space while an unrelated one lands much farther away. It then builds toward semantic search: embedding both a user's question and a database of documents, then retrieving whichever documents' vectors land closest to the question's vector, so a search for "How can I get back into my account?" can surface documentation titled "Recover access to your account" even without matching words. From there, the episode covers retrieval augmented generation, where a system embeds and stores a company's documents, retrieves the most relevant ones for a given question, and hands them to a language model as context. It's direct about the limits, too: different embedding models trained on different data for different tasks can produce different vector spaces and different results, so an embedding is never a perfect stand-in for everything a piece of text means. The episode closes on multimodal embeddings, where Hugging Face's own documentation describes shared embedding spaces that let a text query retrieve related images, extending the same mathematical idea beyond text alone. Sources & References Hugging Face: Using Sentence Transformers — https://huggingface.co/docs/hub/en/sentence-transformers Hugging Face: sentence-transformers/all-MiniLM-L6-v1 — https://huggingface.co/sentence-transformers/all-MiniLM-L6-v1 Hugging Face: Semantic Search with FAISS — https://huggingface.co/docs/course/chapter5/6 Hugging Face: Multimodal Embedding & Reranker Models with Sentence Transformers — https://huggingface.co/blog/multimodal-sentence-transformers Voice narration is AI-generated.

  7. 4d ago

    Quick AI News - 10/02/2026 - Gemini 4 Argon, Anthropic's $42B Broadcom Deal, and AI Becoming Infrastructure

    Today's AI news points to the same thread running through several unrelated stories: AI moving from standalone software into infrastructure, workflows, and regulated systems. Google released Gemini 4 Argon on September 30, its new frontier model built for long-running, complex work in software engineering, legal and financial knowledge work, and cybersecurity, with an industry-leading 1 million output tokens. Access is initially limited to trusted cybersecurity defenders through Google's Fairwind program while the company gathers feedback and builds safeguards before a wider release. Google says it is already using Argon internally for code migrations and cybersecurity work, though those are the company's own claims pending independent use. The biggest business story is financial: Anthropic's IPO filing revealed that Broadcom has agreed to lend Anthropic up to $42 billion to help finance infrastructure spending, tied to Anthropic's commitment to roughly $125.2 billion of TPU computing capacity over five years, per Reuters. The arrangement means Broadcom is acting as chip supplier, lessor, and lender all at once. On the security side, OpenAI says it disrupted a coordinated campaign that attempted to extract its models' protected reasoning, peaking at roughly 16,000 requests from more than 4,000 users on July 24 and 25, which it links to a cluster associated with Moonshot AI, the company behind Kimi. OpenAI says the campaign didn't involve breaching encryption or stored conversations. Separately, OpenAI parted ways with three safety researchers the Wall Street Journal reports were accused of sharing confidential information with an outside AI safety organization. On the consumer side, ChatGPT added a feature letting users virtually try on clothing from a selfie inside shopping results. And on policy, California signed new workplace AI protections restricting employers from relying solely on AI for disciplinary or termination decisions, requiring transparency when AI drives workforce reductions, and limiting the use of AI to interpret workers' emotions. Sources & References Google: Introducing Gemini 4 Argon — https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/ Reuters: Broadcom to lend Anthropic up to $42 billion to lease its chips — https://www.reuters.com/business/broadcom-lend-anthropic-up-42-billion-lease-its-chips-filing-says-2026-10-01/ OpenAI: Disrupting a coordinated model-distillation campaign — https://openai.com/index/disrupting-a-coordinated-model-distillation-campaign/ TechCrunch: OpenAI cuts ties with three safety researchers — https://techcrunch.com/2026/10/01/openai-cuts-ties-with-three-safety-researchers-wsj-reports/ OpenAI Help Center: Shopping with ChatGPT Search — https://help.openai.com/en/articles/11128490-shopping-with-chatgpt-search TechCrunch: ChatGPT can now virtually try on clothes for you — https://techcrunch.com/2026/10/01/chatgpt-can-now-virtually-try-on-clothes-for-you/ California Governor: California's AI workplace protections — https://www.gov.ca.gov/2026/09/30/californias-nation-leading-ai-framework-just-got-stronger-governor-newsom-signs-more-first-in-the-nation-worker-protections-and-more/ Voice narration is AI-generated.

  8. 5d ago

    072 - How to Use Open-Weight Models

    What does it mean to use an open-weight AI model instead of just a chatbot through someone else's service? This episode walks through the practical path: finding a model, checking what you're allowed to do with it, running it, and deciding whether to fine-tune it. An open-weight model makes its trained parameters, or weights, available to download and run, though licenses, training data, and code can each carry their own level of openness, separate from the weights themselves. Developers typically find these models on Hugging Face, a large model and dataset repository where pages include a model card covering intended uses, limitations, evaluations, and licensing, and where some models are gated and require requesting access from the model's authors. Once you have a model, you choose where to run it: a hosted service, a rented cloud GPU, or your own hardware, where Hugging Face Transformers and tools like llama.cpp can load it and run inference. The episode explains GGUF, the model file format llama.cpp commonly uses, and quantization, which represents weights with fewer bits to reduce memory use and run on consumer hardware, with a quality tradeoff depending on how aggressively the model is compressed. When a model runs but doesn't behave the way you want, the episode covers fine-tuning: training a pretrained model further on your own examples using Hugging Face's PEFT library and methods like LoRA, or Low-Rank Adaptation, which trains a small set of additional adapter parameters instead of the entire model, and QLoRA, which combines that approach with quantization to cut memory requirements further. It's direct about the tradeoffs too: fine-tuning isn't the right tool for frequently changing information or simple workflows, dataset quality drives the outcome more than any other factor, and downloading a model's weights doesn't cover the separate licensing terms that apply to commercial use. The practical workflow it lays out: pick a model by capability, size, license, and hardware fit, download it, run it as-is, quantize if needed, and only fine-tune once prompting, retrieval, or tools fall short. Sources & References Hugging Face: Transformers Quickstart — https://huggingface.co/docs/transformers/quicktour Hugging Face: Models — https://huggingface.co/docs/hub/main/models Hugging Face: Model Cards — https://huggingface.co/docs/hub/main/model-cards Hugging Face: Gated Models — https://huggingface.co/docs/hub/models-gated Hugging Face: PEFT Quicktour — https://huggingface.co/docs/peft/quicktour Hugging Face: Parameter-Efficient Fine-Tuning — https://huggingface.co/docs/transformers/peft Hugging Face: LoRA — https://huggingface.co/docs/peft/en/package_reference/lora Hugging Face: Datasets — https://huggingface.co/docs/hub/datasets llama.cpp: GitHub — https://github.com/ggml-org/llama.cpp llama.cpp: Obtaining and Quantizing Models — https://github.com/ggml-org/llama.cpp/blob/master/docs/models.md Voice narration is AI-generated.

About

AI explained in bits. Each episode takes one concept, like tokens, embeddings, hallucinations, or prompt injection, and explains it in about five minutes. No jargon, no filler. Just the idea, why it matters, and what to remember. If you're curious about AI or already building with it, you'll come away understanding how these systems work. One concept. Five minutes. That's the whole show.

You Might Also Like