In this episode of Tank Talks, host Matt Cohen sits down with Keegan McCallum, co-founder and CEO of uRUN, an inference provider built for persistent, steerable, interactive AI experiences. Keegan’s path took him from jailbreaking iPods in Thunder Bay, Ontario, to building the infrastructure that powered Luma AI’s viral Dream Machine launch, scaling from 500 to 9,000 H100s in six hours to process over half a million videos. They break down why the industry is wrong to shoehorn interactive AI into old request-response systems and why the future belongs to persistent, stateful, real-time experiences. Keegan explains the technical differences between LLMs and video models, why infrastructure is the biggest bottleneck for the next wave of AI, and what it takes to recruit top engineering talent. He also shares hard-won lessons on selling to developers, the Canadian AI talent landscape, and his bold, contrary views on the AI supply chain. From avatars and steerable video to coding agents generating 1,000 tokens per second, this episode is packed with frontline insights. Whether you are a founder building in AI, an infrastructure engineer, or simply curious about where human-computer interaction is headed, Keegan delivers a clear-eyed look at the future of interactive AI. – A big thanks to our sponsor, Moomoo Canada This is the kind of tooling that used to live on a Bloomberg terminal, but now it is on your phone, just a few taps away. They offer real-time data, full options chains, and an AI assistant that actually explains trading strategies. Moomoo is the perfect place for people who want to take their money seriously. Open an account today at moomoo.ca The Luma AI Experience: From 500 to 9,000 H100s (04:18) * Joining Luma after building an early LLM infrastructure startup * The Sora demo and the pitch to build and release a competing video model * Building the inference infrastructure for Luma’s launch * Scaling from roughly 500 to 9,000 H100s in six hours as Dream Machine went viral What a Viral AI Launch Actually Looks Like (06:03) * The request queue hitting 100,000 during the Dream Machine launch * Processing half a million videos in 12 hours and reaching one million users in four days * Why having a queue protected the user experience during the surge * Building infrastructure that could run across raw VMs, Slurm, Kubernetes, and different GPU providers * Why founders need to plan for success before the product takes off Why AI Production Infrastructure Is Being Left Behind (10:50) * The tension between training the next model and supporting production workloads * Why infrastructure teams can lose talent to model research * The opportunity Keegan saw in building a dedicated ML infrastructure team * Creating infrastructure that model labs across different modalities can use instead of rebuilding it themselves Why Video AI Needs a Different Infrastructure Stack (12:32) * The limitations of moving media through traditional object storage * The latency problems that make real-time previews and interaction difficult * Streaming audio, video, images, and text as continuous inputs and outputs * uRUN’s focus on making persistent, interactive AI experiences easier to build * Moving AI from something you message toward something that feels present with you The Shift Toward Omnimodal and Interactive AI (19:01) * What Keegan learned working closely with Luma’s research team * Why understanding the research direction matters when building infrastructure * The move toward AI that can understand vision, audio, language, reasoning, and sensor data * How AI has evolved from single prompts to conversations, agents, and interactive systems * Why highly interactive, real-time AI experiences could become the next major interface Small Models, Specialized AI, and the Future of Inference (26:31) * Keegan’s previous startup thesis around smaller models and delegated inference * Why continual learning remains a difficult problem * Running capable smaller models locally and knowing when to delegate to larger models * Why enterprises may increasingly want differentiated models running in-house * The potential shift toward using large models for the hardest problems rather than every task Where uRUN Is Seeing Early Demand (30:38) * AI avatars for customer support and sales * Steerable video generation that lets users intervene while content is being created * Video transformation for creators, avatars, and live experiences * Real-time visual effects for music festivals and live events * Why world models could eventually have applications in open-world gaming The Infrastructure Problem Holding Back World Models (00:32:54) * Companies developing world models without the infrastructure to serve them * Why scaling world models remains difficult * The enormous context requirements involved in tracking everything happening in a simulated environment * The need for distributed infrastructure as these models become more capable Building a Team Around Hard Problems (35:18) * Finding engineers with deep experience in infrastructure and low-latency inference * Bringing together expertise from New Relic, AWS, Superorbital, and other infrastructure environments * Why Keegan looks for engineers who can work across multiple areas * Building a small team of high-leverage people instead of optimizing for headcount The Trait Keegan Looks for in Exceptional Engineers (37:47) * Why culture and quality matter more than quantity * Looking for people who take ownership and solve problems without waiting for instructions * The importance of urgency and experience operating under pressure * Why hard problems can attract customers, investors, and exceptional talent * Hiring people who can wear multiple hats and produce at a high level Building AI Companies in Canada (41:29) * Why Keegan keeps finding Canadians throughout Silicon Valley * The advantages of building a strong team in Canada at a lower cost * Government funding and the opportunity to build ambitious companies with smaller teams * The challenge of Canadian companies and talent being pulled toward Silicon Valley * Why Canada needs more ways for companies to grow instead of selling early Selling Infrastructure to Developers (44:23) * Why developers want to try the product rather than hear a list of performance claims * Giving technical users something they can get their hands on and test * Why open-source surfaces can help developers understand how a product works * Documentation and developer experience as a core part of selling infrastructure * Why developers have a very high tolerance bar for technical errors The Infrastructure Shift Coming Next (46:43) * Why persistent streaming infrastructure could become standard * Moving away from request-response systems toward long-running streams * The importance of maintaining state across continuous interactions * Potential applications in real-time video, virtual try-on, image editing, and coding agents * The infrastructure required for AI systems that can generate and respond almost instantly About Keegan McCallum Keegan McCallum is the co-founder and CEO of uRUN, an inference provider built to handle persistent, steerable, and interactive AI experiences. A self-taught engineer who started by jailbreaking iPods in Thunder Bay, Keegan has built a career on tackling the hardest infrastructure problems in Canada. He has held leadership roles at Colony Networks and served as the Head of Engineering at Luma AI, where he led the infrastructure team that scaled the viral Dream Machine launch from 500 to 9,000 GPUs in a single day. His deep technical expertise spans ML infrastructure, cloud-native systems, and low-latency edge computing. Connect with Keegan McCallum on LinkedIn: https://www.linkedin.com/in/keeganmccallum3?originalSubdomain=ca Visit uRUN’s website: https://urun.sh/ Connect with Matt Cohen on LinkedIn: https://ca.linkedin.com/in/matt-cohen1 Visit the Ripple Ventures website: https://www.rippleventures.com/ This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit tanktalks.substack.com