Leadership for the Physical AI Age

Titto Thomas

A podcast by Tryfecta Capital exploring the intersection of AI, Robotics, and Leadership. We deconstruct the business of Physical AI: how to invest in it, how to build it, and how to lead through it. Featuring interviews with Industrialists, deep tech founders, VC insights, and market analysis on the future of automation and embodied intelligence.

  1. 2d ago

    Ep 13: Beyond LLMs: Why Vision-Language-Action Models Are the Future of Physical AI

    If you are an engineer, this is the one to stay for. Prerna Dhareshwar of Voxel51 returns with Nathan Maroney to go a layer deeper than Episode 12, into what actually runs inside a physical AI system. The short version: these models predict the next action the way a language model predicts the next token, and that one architectural fact quietly dissolved a problem the industry spent years solving in hardware. Prerna works on the product side at Voxel51, where she led the platform that physical AI teams use to explore, visualize, curate and query their data. She came to it from the other end of the problem: an engineering degree from IIT Madras, a Master's from Stanford, research at India's National Aerospace Laboratories, predictive analytics at Pure Storage, and vision based anomaly detection for manufacturing at Instrumental. She has seen this from the model side, the data side, and the factory floor. Start with VLMs. They are the large language models you already know, trained on images and video alongside text, so they carry an understanding of what they are looking at. Then VLA models, Vision-Language-Action. They take every sensor input, take a language prompt describing the goal, and output action tokens continuously, each one conditioned on what the sensors are saying at that instant. Prerna walks it through with a robotic arm unloading a dishwasher. At timestamp zero it sees the dishes and decides its next move is to reach for a plate. That changes the inputs. Now the next action is to grasp. And so on. The model was never trained on your dishwasher, or Nathan's, and it does not need to be. The consequence is the most contested claim in the episode. Because the model conditions on whatever it has at each moment, the sensor streams do not need to be time synchronized. In her words, alignment of sensors is not really something people are too worried about anymore. Then trust. Titto puts the noise problem to her using his own house. A spotless kitchen is one thing. A dish sitting on the roof is another. Real environments are not controlled, and mathematically the difference is just noise. Her answer runs through post-training, the same human alignment step that makes language models sycophantic, applied instead to a human critiquing each action a robot takes. In autonomous driving that is the gap between a car that is safe and a car that behaves the way other drivers expect. Early Waymos followed the road rules exactly and got rear ended by humans who do not. Nathan names the failure mode nobody wants to discuss. Industrial pilots that succeed technically and fail commercially. Heavy industry generates enormous volumes of sensor data, but identifying the small slice worth training on takes a specialist team, and operational leaders have KPIs tied to throughput rather than to technical change. Resistance is not ignorance, it is incentives. The episode closes on Voxel51's platform, why edge cases and the long tail decide the last fraction of a percent, and why Prerna is bullish on physical AI while still calling it early. CHAPTERS 00:00 What is a VLM, and why it matters 01:36 Multimodality and where the models are heading 03:52 Sensor coverage across a mine the size of a city 04:55 Why operational KPIs block adoption 06:29 Selling a model to a board 07:55 What happens when sensors are not time synced 08:31 VLA models: next action prediction explained 09:29 The dishwasher, step by step 11:23 Why sensor fusion stopped being the problem 13:10 The world's first fully autonomous rig 13:39 Noise, uncontrolled environments, and the dish on the roof 16:51 Waymo, road rules, and getting rear ended 17:20 Industry 5.0 and human in the loop 18:12 Post-training, sycophancy, and human alignment 20:20 Manufacturing: from defect detection to assembly 22:25 Pilots that succeed technically and fail commercially 24:44 Inside Voxel51's platform 26:10 Bullish, but early 27:05 Defense, swarms, and a higher bar 30:22 Edge cases, the long tail, and the last 0.9% GUESTS Prerna Dhareshwar, Voxel51 https://www.linkedin.com/in/prernamd/ Nathan Maroney, Director, Tryfecta Group Host: Titto Thomas, Managing Partner and Co-Founder, Tryfecta Group LINKS Watch on YouTube: https://youtu.be/AFuHb-vLbyo Voxel51: https://voxel51.com Tryfecta: https://tryfecta.biz

    Ep 13: Beyond LLMs: Why Vision-Language-Action Models Are the Future of Physical AI
  2. Aug 24

    Ep 12: Beyond Vision: Why Multimodal Data and Software Power Physical AI

    Computer vision has been around for decades. So why is physical AI suddenly everywhere, with a new humanoid robotics company funded almost every week? Prerna Dhareshwar of Voxel51 and Nathan Maroney of Tryfecta Group join Titto Thomas to explain what changed, why vision alone will never be enough, and why the real differentiator is the software and data layer that almost nobody is talking about. Prerna traces the shift back to what large language models taught the field about generalization. Language modeling is next token prediction. Physical AI is next action prediction. Once it became clear the same architectures could work in both, capital and attention followed. Prerna is on the product team at Voxel51, where she led a product that helps physical AI teams explore, visualize, curate, and query their data. Before that she was a machine learning engineer building vision based anomaly detection for manufacturing at Instrumental. Nathan Maroney, Director at Tryfecta Group and the group's mining lead, brings the operator view from mining and heavy industry, where you are processing hundreds of thousands of tons and a small consistency gain in concentration is worth real money. The conversation gets practical fast. Nathan asks the question every operator is actually asking: how does a mine site get itself ready for this, and how is physical AI different from the conventional automation, machine vision, and predictive maintenance they already have? Prerna's answer is to look sideways. Auto manufacturing OEMs are already using humanoid robots to build and assemble parts, in exactly the complex three dimensional work that robotic arms could never automate away. Find what worked in an adjacent industry, then lift and generalize it. That approach also de risks the first move. On whether physical AI is just self driving cars, Prerna points out that Waymo has had more than a ten year head start collecting data, which is why that sector looks a few steps ahead of everyone else rather than being the whole story. There is an honest detour into risk. In San Francisco, Prerna keeps seeing ads for humanoid robots that will clean your house for a flat fee of $150 regardless of size. Great deal, until it leaves you with a pile of broken dishes. Titto's line for how heavy industry thinks about that: it is not the early bird that gets the worm, it is the second mouse that gets the cheese. Prerna's view is that the calculus has genuinely changed, and that she would have answered differently a year or two ago. Then the technical core. Physical AI data is hard because the sensors are not time synchronized. Each one records at its own frame rate and frequency, across long time ranges, and all of it has to be aligned, curated, labeled, and fed downstream before a model ever sees it. That tooling problem is where a lot of the real work lives. On why multimodal beats vision only, Prerna uses the human analogy. We do not perceive the world through sight alone. Two eyes give us depth. Touch tells us how much pressure a delicate object can take. Restrict a system to a single camera and you have severely limited what it can know about the world it is operating in. She takes a position on the vision only versus lidar debate, citing an edge case where a Waymo could not see a person crossing from behind a parked truck, and the lidar caught what the cameras could not. Titto brings a story from the Apache gunship program, where the sensor lens was cut from pure sapphire because the software of the 1980s could not correct for chromatic aberration. Everything had to be fixed at the physical layer to hand the software a clean image. Today a ten dollar sensor does the same job, which is exactly the hardware to software shift the field is living through. With VLMs, the sensors are fixed and the goal is fixed as a text prompt. What the system controls is the sequence of actions it takes to get there. Nathan closes on the practical blocker. Simulating a controlled separation process is achievable low hanging fruit, and most of the sensors are already installed, but a site typically has to wait two years to accumulate enough data before it can deploy. He wants systems that are self learning, self healing, and self governing from day one on a greenfield site. Prerna's answer is that it is never too early to start collecting data, and that even if you are years away from deploying anything, making sure it is clean, structured, and stored properly is the move available to you right now. She connects the self learning ambition to meta learning and few shot generalization, and to what today's models already do in a limited way through context. Prerna is coming back for a deeper technical episode on VLMs and how to deal with noise. CHAPTERS 00:00 Welcome and introducing Prerna Dhareshwar01:06 Computer vision is not new, so what actually changed01:22 Next token prediction to next action prediction03:55 Nathan on mining: where the margin really sits04:57 Lifting proven applications from adjacent industries07:37 Is physical AI just self driving cars09:08 The $150 humanoid and the broken dishes problem10:33 Are we there yet, and how much should we fear it11:56 Convincing conservative operators to move13:46 Why software is the real differentiator15:53 Multimodal versus vision alone17:51 Vision only or lidar, and the Waymo edge case18:55 Trust, and the Apache gunship sapphire lens21:09 From hardware to software: what VLMs changed23:10 Process simulation and the two year data wait25:31 Start collecting data now, and the path to self learning28:19 Next time: VLMs and filtering out noise GUESTS Prerna Dhareshwar, Product, Voxel51https://www.linkedin.com/in/prernamd/ Nathan Maroney, Director, Tryfecta Group Host: Titto Thomas, Managing Partner and Co-Founder, Tryfecta Group LINKS Watch on YouTube: https://www.youtube.com/watch?v=o7Rx3sNpgWMVoxel51: https://voxel51.comTryfecta: https://tryfecta.biz

    Ep 12: Beyond Vision: Why Multimodal Data and Software Power Physical AI
  3. Aug 2

    Episode 11: From Mission Control to Machine Control: Inside the Physical AI Stack with Sami Sultan

    Automation does not fail at the top. It fails in the layers underneath. Sami Sultan, Vice President at Darcy Partners, ex BCG, and ex Shell wells engineer, returns for part two: a walk through the full Physical AI stack as it actually exists in heavy industry today. In this episode: The stack, layer by layer. Cameras reading rock as it comes off the shale shaker, computer vision flagging anomalies, AI agents that sense, plan, and act, and the actuation layer where a digital decision finally touches physical equipment. Human in the loop. Why the geosteer keeps their job when a few percent of production is worth millions, and why safety critical calls will stay human for a long time yet. The strange origin of mission control. How a failed military operation in 1980s Iran gave birth to joint command, then to the remote operations center, now the brain of rigs, mines, and factories everywhere. Why full stack automation ventures go bankrupt. One sensor fails and the house of cards comes down. The aerospace lesson is that automation is a management system, not a feature. Beyond the buzzword. Why Sami refuses to say digital twin, and why physics informed neural networks are winning the trust that black box AI cannot. And the next frontier. Drilling techniques crossing into mining, continuous extraction, and the bridge that could finally make Western rare earth production viable. Watch this space. About the guest: Sami Sultan is Vice President of Oil and Gas and AI at Darcy Partners, where he leads technology scouting and advisory for the world's largest energy operators and utilities. He spent nearly a decade at Shell as a wells engineer, where he helped found Shell Geodesic, an algorithmic well navigation venture that applied AI to the subsurface years before it was mainstream. He holds seven patents, is a BCG alum, and earned his MBA in sustainability from the Yale School of Management.

    Episode 11: From Mission Control to Machine Control: Inside the Physical AI Stack with Sami Sultan
  4. Jul 20

    Ep 10: Powering the AI Boom: Distributed Energy, Microgrids, and the Future of Oil & Gas with Sami Sultan

    The AI boom has a power problem, and simply scaling up traditional energy infrastructure will not solve it. Sami Sultan, Vice President at Darcy Partners, ex Shell wells engineer, and BCG alum, joins Episode 10 to map where the energy for the machine age could actually come from: microgrids, distributed energy resources, and a pragmatic transition where existing capabilities get repurposed rather than expanded. In this episode: The power question: why energy demand compounds as AI spreads, and why the sustainable path runs through distributed generation, not just bigger grids. The transition in practice: stranded gas bridging data center demand today, old wells repurposed for energy storage tomorrow, and how capital decides what gets built. What heavy industry brings to the table: decades of experience moving liquids, running remote operations, and managing complex infrastructure, now pointed at new problems. The Physical AI stack arriving on sites today: red zone cameras, emissions drones, robot dogs, and the climb from sensing to agents to autonomy. Sami's origin story: building algorithmic well navigation at Shell in 2017, trained on synthetic wells the way Waymo trained on synthetic miles, before most boardrooms knew what a GPU was. And a first look at the three layer architecture we are building at Tryfecta: an agentic base, a physics model at the core, and command and control on top. Part one of two. Sami returns next episode for the technology deep dive.

    Ep 10: Powering the AI Boom: Distributed Energy, Microgrids, and the Future of Oil & Gas with Sami Sultan
  5. Jul 13

    Ep 9: Cutting Rates with Robots: Capital Flows and the Deflationary Power of Physical AI

    What if the fastest way to cut interest rates for the whole world is to teach machines to mine? Daniel Dangoor (Investments and Treasury) and Nick Shelton return for Episode 9, and this time we follow the money. Capital has been pouring into the poster children of Physical AI, drones, humanoid robots, and driverless cars, while the real prize sits underneath: machines that sense, extract, and build in the physical economy. In this episode: The deflation thesis: when Physical AI cuts the cost of mining and energy, supply rises, commodity prices fall, and the world gets easing no central bank can deliver. The iPhone economy already proved the mechanism. Why the middle of the commodity supply chain gets crushed in every cycle, and what junior miners teach us about survival. The hyperscaling question: software was the one sector that could scale 100x, and AI just ended that monopoly. Where do outsized returns come from in a physical world? Whoever has more robots wins: the case for effectively infinite capital flowing into robotics, and why debt that builds GDP is not the problem people think it is. Industry 3.0 to 6.0: from the space race that created Intel to the coming era where machines lead. The people side: why Silicon Valley is hiring problem solvers, because nobody can define an AI engineer yet. Dan closes with the best analogy of the series so far: when a person loses one sense, the others sharpen. When humanity hands its base skills to machines, watch what the remaining ones do.

    Ep 9: Cutting Rates with Robots: Capital Flows and the Deflationary Power of Physical AI
  6. Jul 5

    Ep 8: The Compute Space Race: Geopolitics, US Hegemony, and the Physical AI Gold Rush

    How does a baby learn faster than an LLM? Not by reading more text, but by touching the world. That analogy from Daniel Dangoor (Investments and Treasury) anchors this episode's thesis: language models are capped by the finite supply of human text, and the next leap in AI depends on machines that can sense the physical world. Host Titto Thomas, Daniel Dangoor, and Nick Shelton unpack why sensors, not chatbots or humanoid robots, are the underserved gold rush of Physical AI. In this episode: Physical AI is bigger than humanoid robots and driverless cars. From rig sensors at Shell that optimized an entire fleet, to in situ soil analysis that maps rare earth deposits in a day instead of 3 months. The compute space race. Dan's macro thesis on why the US treats AI as a race it must win at any cost, and why that makes the compute investment supercycle effectively unlimited. Why sensors are the new Nvidia trade. Sensor stocks lagged every AI basket for 18 months, then rallied 80% between April and June 2026 as real industrial demand, not speculative hype, finally arrived. The ethical scaffolding. Drawing on their backgrounds in theology and philosophy, the panel asks whether governance is mature enough for the productivity and geopolitical stress ahead. Solving humanity's dirty jobs. Why machines should handle the 12 hour pipe inspections in the desert so people never have to. The takeaway: language is only the beginning. The industrial economy needs AI that can feel, and capital is now shifting to build the sensors that make that fusion possible.

    Ep 8: The Compute Space Race: Geopolitics, US Hegemony, and the Physical AI Gold Rush
  7. Jun 29

    Episode 7: Harnessing the Winds of Physical AI: The Change Management Blueprint for a Human-Centric Future

    While the vision of Physical AI is incredibly optimistic, the actual implementation is going to be messy. As AI native companies see productivity skyrocket, they are also facing a hidden crisis: a massive drop in workplace empathy as managers get used to bossing around AI agents 24/7 and expect the same from their human employees. In this episode of Leadership for the Physical AI Age, Nick Shelton returns to discuss the critical importance of change management as we enter the next industrial revolution. Host Titto Thomas breaks down Tryfecta's approach to building a "tag team" between human workers and AI agents, and why the ultimate human advantage will always be first-principles critical thinking. Using the tragic lesson of a 1999 Swissair crash, Titto explains why rigid, checklist-based training is dead, and how human common sense must step in to oversee the AI processes of the future. Finally, Nick and Titto revisit the ultimate question: How do we use this fire to cook our food, instead of burning down our house? In this episode, we cover: The Empathy Deficit: Why interacting with 24/7 AI agents is changing human behavior and creating friction in the workplace.The Tag-Team Model: How Tryfecta is designing systems where agents handle the heavy lifting, but call in humans for critical "last mile" process safety.The 5-Year Skill Shelf Life: Why continuous learning and education funds are the only way to survive the modern war for AI talent.The Swissair Lesson: How rigid checklists fail in crises, and why first-principles thinking is the most important skill to teach the next generation.The Good Life: Using AI to automate mundane administrative tasks to reclaim time for family, legacy building, and deep work.Learn more and connect with us: Visit our website: tryfecta.bizFollow Titto Thomas on LinkedIn: Titto ThomasFollow Nathan Maroney on LinkedIn: Nathan MaroneyFollow Nick Shelton on LinkedIn: Nick Shelton

    Episode 7: Harnessing the Winds of Physical AI: The Change Management Blueprint for a Human-Centric Future
  8. Jun 22

    Ep 6: The Diamond Workforce: Hacking the Future of Jobs and AI Agents

    Are we all going to lose our jobs to AI? As the AI revolution accelerates from digital clouds to physical industries, anxiety around the future of work has never been higher. In this episode of Leadership for the Physical AI Age, Tryfecta Capital host Titto Thomas sits down with Nick Shelton—an early Google veteran, former recruiter, and startup scaling expert who helped build a massive $15B autonomy powerhouse. Nick brings a deeply human perspective to the AI conversation, arguing that the traditional "pyramid" organizational chart of the Industrial Revolution is about to become a "diamond." As AI agents take over entry-level data tasks, the future belongs to those who learn to manage and build alongside digital twins. Titto and Nick discuss the shift from hardware dominance to software supremacy, how first-principles thinking is rewriting the startup playbook, and the exact skills the next generation must learn to thrive alongside Physical AI. In this episode, we cover: The Google Blueprint: Nick's journey from early Google sales to scaling billion-dollar autonomy startups.The Diamond Workforce: Why the traditional corporate pyramid is flattening, and how AI agents will serve as the new entry-level workforce.The Hardware-to-Software Shift: How a $10 camera today replaces a $3 million hardware canopy from the 1980s.Upskilling for the AI Age: Practical advice for young professionals and seasoned managers on adapting to a world of AI agents and digital twins.The Future of Teams: Why the next generation of billion-dollar companies will be run by incredibly small, hyper-focused teams.Learn more and connect with us: Visit our website: tryfecta.bizFollow Titto Thomas on LinkedIn: Titto ThomasFollow Nathan Maroney on LinkedIn: Nathan MaroneyFollow Nick Shelton on LinkedIn: Nick Shelton

    Ep 6: The Diamond Workforce: Hacking the Future of Jobs and AI Agents

About

A podcast by Tryfecta Capital exploring the intersection of AI, Robotics, and Leadership. We deconstruct the business of Physical AI: how to invest in it, how to build it, and how to lead through it. Featuring interviews with Industrialists, deep tech founders, VC insights, and market analysis on the future of automation and embodied intelligence.