NVIDIA Generative AI

Cloudadorn Academy

What is actually inside the phrase generative AI, and why does every answer end up being a cost decision? NVIDIA Generative AI is a twenty-three-episode series from Cloudadorn Academy for people who want the real picture and are not already specialists. Shelley asks the questions a curious adult would actually ask. Rob answers with something you can picture first and the name of the idea second: nesting dolls for the AI hierarchy, a group chat for attention, sand and statues for diffusion, a case conference for sensor fusion. Season one walks four acts. What these things are, from the hierarchy and the learning paradigms through attention, training, and the architecture zoo. How you make a model yours for the least money that works, from prompting to retrieval to LoRA to a full fine-tune, and which number tells you it worked. How a model gets eyes and ears, through shared embedding spaces, patches as tokens, diffusion, segmentation, and sensor fusion. Then the stack itself: precision formats and memory bandwidth, serving and packaging, the training pipeline, agents, guardrails, robots and world models, and why the ecosystem is hard to leave. NVIDIA is the through-line because it sells at every one of those layers, so its product map doubles as a map of the field. Where a vendor number cannot be confirmed against a primary source, the hosts hedge it or drop it and say which one they did. This series is narrated using AI voice technology. The content and scripts are original. Shelley and Rob are original hosts, not impersonations of real people. If the field finally holds still long enough to make sense, subscribe, leave a review where you listen, and visit cloudadorn.com.

  1. Épisode 1

    S1E01. Nesting Dolls: What Generative AI Actually Contains

    This episode is narrated using AI voice technology. The content and script are original. Artificial intelligence, machine learning, deep learning, generative AI. Four phrases, used interchangeably, by the same person, about the same product. They are not synonyms. They are nested, and every step inward is a stronger claim about what is actually happening. Shelley makes Rob open the dolls one at a time, biggest to smallest. You will learn why the nesting only runs one way, so a chess program built from hand-written rules is AI and is not machine learning. Why foundation model is not a fifth doll at all: it describes a model trained wide enough to be reused, which is not the same as a model that makes things, and CLIP is the counterexample that proves it — broad, reusable, and it makes nothing at all. NVIDIA's Cosmos family is the other half of the point: world foundation models that generate video of the physical world, also not language models. The difference between learning the boundary between categories and learning the shape of the data, which is the whole reason a model can make something new instead of only sorting what exists. And the four ways a model learns, including the one people get backwards: pretraining a large language model is self-supervised, not unsupervised, because hiding the next word turns the sentence into its own answer key. Then the arc. 2012, when a deep network won the ImageNet competition by a distance. 2017 and the paper that threw out recurrence. 2020 and a model doing new tasks from examples in the prompt. 2021, when the same architecture came for pictures. Underneath all of it, multiplying big grids of numbers on a chip that was built for video games. Takeaway: when a product says AI, it has told you almost nothing. Ask which doll. Subscribe for the rest of the season, and visit cloudadorn.com.

  2. Épisode 2

    S1E02. Attention: How a Model Reads a Sentence

    This episode is narrated using AI voice technology. The content and script are original. You type a sentence. Something types back. This is what happens to your words in between, in order, with the clever bit slowed right down. Rob walks Shelley through the five steps every one of these models runs: the sentence is chopped into pieces smaller than words, each piece becomes a location in a space, position gets added on purpose, the stack of blocks runs, and one word comes out. Then it all runs again. The step that sounds like housekeeping turns out to be the one worth stopping on: attention is blind to order. It sees a heap, not a line. Leave that step out and "the dog bit the man" and "the man bit the dog" are the same input to the model. Then attention itself, which is smaller than its reputation. Every token asks a question, advertises what it knows, and offers something to contribute. Everything gets weighed against everything else, and nothing is ever ignored, only weighted near zero. That last part comes with a bill: ten pieces is a hundred comparisons, twenty is four hundred. Double the length, quadruple the work. That single curve is why long context is expensive, and every trick in the back half of this episode is somebody trying to get out from under it. Also: why a blindfold is the reason text models are built the way they are, how to tell a model's shape from the job it does, and why one model honestly has two parameter counts ten times apart. NVIDIA's Nemotron 3 Nano has 31.6 billion parameters and runs 3.2 billion of them per token. When somebody quotes you a parameter count, ask which one.

À propos

What is actually inside the phrase generative AI, and why does every answer end up being a cost decision? NVIDIA Generative AI is a twenty-three-episode series from Cloudadorn Academy for people who want the real picture and are not already specialists. Shelley asks the questions a curious adult would actually ask. Rob answers with something you can picture first and the name of the idea second: nesting dolls for the AI hierarchy, a group chat for attention, sand and statues for diffusion, a case conference for sensor fusion. Season one walks four acts. What these things are, from the hierarchy and the learning paradigms through attention, training, and the architecture zoo. How you make a model yours for the least money that works, from prompting to retrieval to LoRA to a full fine-tune, and which number tells you it worked. How a model gets eyes and ears, through shared embedding spaces, patches as tokens, diffusion, segmentation, and sensor fusion. Then the stack itself: precision formats and memory bandwidth, serving and packaging, the training pipeline, agents, guardrails, robots and world models, and why the ecosystem is hard to leave. NVIDIA is the through-line because it sells at every one of those layers, so its product map doubles as a map of the field. Where a vendor number cannot be confirmed against a primary source, the hosts hedge it or drop it and say which one they did. This series is narrated using AI voice technology. The content and scripts are original. Shelley and Rob are original hosts, not impersonations of real people. If the field finally holds still long enough to make sense, subscribe, leave a review where you listen, and visit cloudadorn.com.