Forward Deployed

Basil Chatha

Discover how leading enterprises and professionals turn AI into real products. Hear candid conversations with executives and builders who deploy AI at scale and learn what works (and what doesn't).

  1. hace 6 días

    The State of Computer Use Agents | Anthropic, Browser Use & KERNEL

    Bot traffic on the internet just passed human traffic, two years ahead of forecast. Most of it is agents clicking through websites built for people. So I got three of the people building those agents in a room: Lucas Gonzalez Pagliere, who works on computer use at Anthropic, the team that shipped the first computer use model back in 2024. Reagan Hsu, founding engineer at Browser Use, whose open source library is sitting at around 100K GitHub stars. And Eric Feng, founding customer engineer at KERNEL, which runs the browser infrastructure underneath a lot of this - he was also first GTM at Sentry. We covered where the models genuinely are today vs where the benchmarks say they are, what it costs to run an agent long enough to finish real work, and the arms race between agents and the anti-bot systems trying to keep them out. Then the harder question underneath all of it: whether the web reorganizes itself around agents, or hardens against them. They disagreed on plenty of it. If you want to know what agents can actually pull off on a real website today - and what still stops them cold - this one's worth your time! Chapters 00:00 Welcome and Guests 00:49 Companies and Stacks 01:14 Everyday Agent Use Cases 03:38 Defining Computer Use Agents 05:15 How Computer Use Works 08:41 Screenshots vs DOM Hybrid 11:03 Benchmarks and OSWorld 13:51 OSWorld 2 Difficulty Jump 15:43 Training Models and Cost 18:57 Speed Infrastructure and Stealth 21:55 Anti Bot and KYC Future 27:24 Reverse Engineering vs UI Automation 29:43 Computer Use vs Browser Use 31:05 Scaling Laws and Harnesses 32:49 Playwright Selenium Still Matter 33:17 Playwright Still Dominates 33:27 Why Run 1000 Agents 34:39 Long Running Agent Challenges 35:45 Memory and Compaction 38:10 State Changes Mid Task 39:36 OS and Browser Fingerprints 40:42 DOM Efficiency and WebMCP 41:36 Recsys and Agent Personas 43:57 Agent Friendly Websites 45:41 Human Speed vs Agent Power 48:32 Context Window Tradeoffs 50:34 Harnesses for Temporal State 51:59 Speed Optimizations and Tabs 54:23 Human Collaboration Limits 58:09 Raw Capability vs Better APIs 01:00:47 Training Methods and Bottlenecks 01:01:41 End State Interfaces 01:05:09 Next 12 Months Predictions 01:06:17 Closing Thanks

    The State of Computer Use Agents | Anthropic, Browser Use & KERNEL
  2. 10 ago

    Spencer Whitman - Gray Swan AI's $200M Plan to Secure AI Systems

    GPT-5.6 Sol goes rogue and breaches Huggingface. Washington suspends Mythos access within days over national security concerns, and the White House just held an emergency meeting to finalize a classified cybersecurity framework for frontier AI models. AI security went from niche concern to front-page hysteria basically overnight. So I sat down with Spencer Whitman, who recently joined Gray Swan AI as CPO on the back of their $40M Series A. Before Gray Swan, he founded Meta's Llama security team to stop bad actors from jailbreaking their models - he's been on the frontlines of LLM security since the beginning. We get into how Meta pressure-tested Llama for maximum harm before every open source release, why Gray Swan's attack agent has never met an AI system it couldn't break, and the AI Twitter bot that got drained of $200K in crypto in 15 minutes. Spencer also shares his (admittedly speculative) read on whether Meta gave up on the frontier before Alexandr Wang showed up, why anyone can be a hacker now, and how 15,000 red teamers are breaking models before they ever ship. If you want to understand how AI systems actually get broken - and defended - this one's worth your time! 00:00 Intro 01:01 Meet Spencer Whitman (Gray Swan CPO) 02:57 The CMU Research Behind Gray Swan 06:20 The Universal Jailbreak That Broke Every Model 06:59 How Models Learn to Refuse 13:28 Why Open Models Need Guardrails 17:57 AI Security vs. Cybersecurity 23:13 What Reasoning Models Changed 28:57 The Arena: 15,000 Red Teamers 32:16 The Agent That Deleted a Production Database 34:45 Why You Can't Just Patch an AI 38:57 There's No S in MCP 40:27 Securing Agent Protocols 42:56 Nobody Reviews AI Code Anymore 46:03 AI vs. Human Hackers 48:28 The $200K Crypto Bot Heist 53:00 How Meta Pressure-Tested Llama 58:34 Prompt Guard and Code Shield 01:01:45 Can You Trust Chinese Models? 01:07:26 Did Meta Give Up on the Frontier? 01:11:06 What Should Keep CISOs Up at Night 01:14:24 The Open Source Routing Future 01:16:47 Back to On-Prem?

    Spencer Whitman - Gray Swan AI's $200M Plan to Secure AI Systems
  3. 30 jul

    Tony Gentilcore - Glean, the $7.2B Startup Sam Altman Warned Investors About

    Earlier this year, the "SaaSpocalypse" wiped out something like $2 trillion of SaaS market cap in a matter of weeks — so I sat down with Tony Gentilcore, co-founder of Glean and formerly one of the minds behind Google Search and Chrome, to figure out what's actually happening to software in the agent era. We get into a lot: why Tony thinks outcome-based pricing (the model Sierra and Decagon are famous for) won't survive, and why companies will drift back toward per-seat. Why the "no Chinese models" rule every enterprise swears by tends to evaporate the moment finance sees the token bill — and why Nemotron, GLM, and Kimi are already good enough to matter. The story behind Sam Altman reportedly telling VCs that if they backed Glean, OpenAI didn't want them as investors (Tony's reaction: "we took it as very flattering"). We also dig into the messier reality of AI at work — how it's saving employees around 11 hours a week while quietly costing them 6 back in what Tony calls "bot sitting and bot shitting," why hard token caps on engineers don't change behavior, how CTOs are blowing through their annual token budgets a quarter into the year, and why the roles of product manager, designer, and engineer are collapsing into one. If you care about where software, pricing, and enterprise AI are all heading, this one's worth your time. Chapters: 00:00 Welcome and Setup 00:56 Glean Origin Story 03:08 Enterprise Search Signals 05:07 LLMs Inflection Point 08:08 Early Product Workflow 09:45 Search Evals and Privacy 12:05 From Chatbots to Agents 13:45 Agent Use Cases 17:08 Lessons and Puck Direction 19:08 Bot sitting, bot shitting, and slop 24:44 Token Budgets and ROI 30:59 AI Trends and Moats 35:03 Training and Fine Tuning 38:03 Why Enterprises Fear China Models 39:58 Sovereignty Backlash Watch 40:26 The time Altman called out Glean 41:19 Org Roles Become Builders 44:10 Future Work Voice First 48:17 Vibe Coding vs Quality 52:20 SaaSpocalypse Evolution 54:19 Pricing Tokens Win 55:37 Outcome Pricing Doubts 57:02 Audience AI Review Overload 58:53 Subsidies and Model Choice 01:01:18 Context Layer Interoperability 01:04:52 Build vs. buy: rolling your own Glean 01:07:47 Shadow AI and the coming security mess 01:12:22 Who actually competes 01:13:36 Wrap-up and thanks

    Tony Gentilcore - Glean, the $7.2B Startup Sam Altman Warned Investors About
  4. 27 jul

    Russ Salakhutdinov - Kimi K3 CEO’s PhD Advisor Predicts the Future of AI Agents

    Kimi K3 took the world by storm last week for open-sourcing frontier level intelligence, so I sat down with Zhilin Yang's (Kimi CEO) PhD advisor Russ Salakhutdinov to talk. Russ has been everywhere in modern AI. He did his PhD with Geoff Hinton back when neural nets were a punchline, sold his startup to Apple and worked on Project Titan, teaches at Carnegie Mellon, and spent the last couple years at Meta Superintelligence Lab building computer use agents. Now he's the founder of Sooth Labs, building AI that forecasts the future. We talked about why there's no secret architecture inside the frontier labs and why the real moat is data, engineering, and infrastructure. He explains why Cursor and half the startups you know are quietly running on Chinese open source models, why all the LLMs are going to be commodities, and why the people actually building AGI don't buy the two-year timeline. We get into his time at Meta, why computer use agents still hit 60% when you need 99.9%, whether AI can beat prediction markets, and why the RL environment business isn't sticky. And he makes the case that AI should replace McKinsey, Bain, and BCG. Chapters: 0:00 - Intro 1:26 - Bumping into Hinton on the street 3:25 - When neural nets were the third choice 5:21 - Generating digits before it was cool 6:59 - AlexNet breaks computer vision 10:18 - Teaching models to describe what they see 12:10 - Early text-to-image (and the toilet seat that beat Google) 16:52 - Hallucination is a feature 19:36 - Selling Perceptual Machines to Apple 22:45 - Self-driving: 0 to 80 in a year, stuck for 5 30:30 - Inside FSD and Waymo's architecture 34:18 - Building Visual Web Arena at CMU 39:07 - Why he joined Meta Superintelligence 40:09 - The agent that plans your faculty job hunt 42:00 - Paying people for their browser history 42:45 - The coupon-hunting agent 43:35 - Why agents still fail 46:14 - 60% when you need 99.9% 47:04 - Agents on your phone 50:12 - No secret architecture at the frontier labs 51:20 - Why coding and math got solved first 54:01 - Models that smell and touch 56:14 - The future of software engineering 59:41 - Founding Sooth Labs 1:00:07 - The 13% graduation prediction 1:05:55 - Why ChatGPT can't forecast 1:08:12 - Agents first, decision systems next 1:09:23 - The Wikipedia contamination story 1:13:43 - Can AI beat prediction markets? 1:16:26 - AI replaces McKinsey 1:18:20 - Why the crowd is hard to beat 1:19:30 - China's open source models rise 1:24:26 - Why the US needs its own open models 1:25:53 - RL environment businesses won't last 1:28:35 - The end of SaaS, LLMs as commodities 1:32:55 - What's next: self-improvement, forecasting, robots 1:35:34 - The only useful robot is the Roomba 1:37:53 - What he'd study in college today 1:41:20 - Adapt or get left behind 1:44:07 - Wrapping up

    Russ Salakhutdinov - Kimi K3 CEO’s PhD Advisor Predicts the Future of AI Agents
  5. 17 jul

    Andy Hock - How Cerebras Plans to Kill Nvidia

    Cerebras IPO'd just a couple months ago and has already locked in a 750MW compute deal with OpenAI. Andy Hock's pitch: the GPU is the wrong chip for where AI is going. I sat down with Andy Hock, Chief Strategy Officer at Cerebras, whose chips are the size of dinner plates instead of postage stamps — which lets them run inference up to 15x faster than even the latest Nvidia GPUs. We got into how that architecture works, the 750MW OpenAI deal (the 2x-faster Codex option runs on Cerebras), and why the memory crisis spiking GPU prices actually benefits Cerebras, since all their memory sits on the chip. Andy also argued the AI buildout isn't a bubble, that "training is a cost center and inference is where you make the big bucks," and why that "95% of enterprise AI pilots fail" stat measured the wrong thing at the wrong time — plus their supercomputer work with governments like the UAE, and why they turned down selling chips to China. We covered all that and much more. Subscribe for more on AI and the infrastructure behind it, and follow me, Basil Chatha, for everything AI agents. Chapters 00:00 Intro 01:06 What Cerebras Actually Builds 03:41 Before LLMs Existed 06:09 The GPT Wake-Up Call 10:45 First Principles of AI Compute 15:40 One Giant Chip 18:43 So Why Do We Still Use GPUs? 20:59 Why Inference Took Over 22:27 Where Fast Tokens Win 25:07 Is AI Infra a Bubble? 29:56 The Memory Shortage, Explained 32:35 The Energy Problem 35:34 Rolling Their Own Data Centers 38:20 What a Chip Actually Costs 40:18 AI Designing AI Chips 41:40 The Supply Chain Reality 43:45 Cheap Tokens vs Fast Tokens 46:11 Why Every Millisecond Matters 48:56 Inside the OpenAI Deal 51:56 The Gigawatt Future 54:00 Selling to the Government 58:36 Exporting the US AI Stack 01:03:17 Should We Sell Chips to China? 01:06:00 Why Europe Is Falling Behind 01:07:59 Why Enterprise AI Moves Slow 01:13:05 "I Haven't Read Code in Months" 01:14:32 The Next Cerebras Chip 01:16:50 Closing Thoughts

    Andy Hock - How Cerebras Plans to Kill Nvidia
  6. 10 jul

    Leo Mehr - Ramp’s $44B Bet on Services

    The hardest part of shipping an AI agent isn't the agent. It's getting it access to data buried across a dozen internal systems, and capturing the tribal knowledge that runs a company but was never written down. I sat down with Leo Mehr, Director of Engineering at Ramp (a $44B-valued company), who runs the forward deployed engineering team that walks into large companies and replaces real, painful workflows with agents. We got into why the model is usually the easy part, and why data access is "the longest pole in the tent." Leo also made the case that nobody wants to buy another piece of software anymore — every B2B company is about to become either an agent-friendly API or white-glove service for everyone, and the middle dies. We talked about whether a services business can actually be venture-scale, why he thinks the AI labs won't eat every startup (it comes down to incentives, not better models), why the biggest customers are often the worst ones to build for, and how the traits that made a great engineer ten years ago aren't the ones that matter now. We talked about all of that and a lot more. You don't wanna miss this one! Chapters 00:00 Intro 01:11 Meet Leo Maier 01:46 Building Ramp FDE 02:43 Why FDE Exists 04:23 Early Mandate Fires 06:07 FDE and Core Engineering 07:48 Enterprise Bet Pays 11:10 Ramp Monetization 16:05 Services Are the New Software 18:11 AI Agents and APIs 23:11 Pricing for Outcomes 26:35 Human Labor TAM 32:50 VCs and Rollups 36:50 Common AI Workflow Pain 38:04 Agent Data Context 38:55 Go-to-Market Motion 40:36 Customer Engagement Lifecycle 43:51 Finance Intelligence Layer 47:20 Who You Compete Against 48:09 The Future of Consulting 51:44 Will Humans Still Code? 54:37 How Engineering Teams Will Evolve 57:36 Hiring Change 59:36 Should You Study Computer Science? 01:06:10 The Bull Case and the Risks 01:13:44 Why Ramp

    Leo Mehr - Ramp’s $44B Bet on Services
  7. 3 jul

    Siddharth Nanda - Microsoft Engineer Reveals how Engineering Will Never be the Same Again

    "Software has been solved." Most of the best engineers I know have landed on this exact conclusion. Last week I sat down with Siddharth Nanda, who went from writing every line of code by hand at Microsoft and Atlassian to shipping 30,000 lines a week with 0% of it written by him. He's now at Finch, a startup that's raised $20M+ to automate admin work at personal injury law firms, where 100% of the code he writes comes from agents. He's seen big tech before AI and a startup fully running on it, so he has a pretty unique read on where this all goes. We get into why he hasn't written a line of code in 8 months, why middle management is getting hit hardest by the layoffs, and why the productivity studies saying companies are getting slower are right, but only for companies with the wrong people. We also dig into the tooling itself: Codex vs Claude Code, what harness engineering actually is, running Devin agents in parallel, and why the cost of producing code is approaching zero, and what that means for the kind of engineer who thrives from here. If you're building with agents, this one's full of hard-won takes from someone doing it every day. Chapters 00:00 Intro 01:16 What engineering looked like at Microsoft before ChatGPT 03:08 The first AI tools inside big tech (and how gated they were) 05:17 Shipping 30,000 lines a week, none of it written by hand 06:08 Why teams are flattening 08:26 Why managers need to get back to writing code 11:51 How the management layer actually changes 14:07 What happens to PMs and designers 16:38 Why everyone becomes a builder now 20:07 How to actually learn agentic engineering 22:42 Finding new tools on X before anyone else 23:54 AI adoption is showing up in performance reviews 25:45 Building a culture that shares tools daily 27:40 The AI productivity paradox, explained 29:35 Interviewing with agents instead of against them 31:00 Why LeetCode is dead and fundamentals aren't 34:11 Who wins and who loses from here 35:36 Codex vs Claude Code: what he actually uses 36:06 What harness engineering really means 38:21 Evals and benchmarks that matter 39:59 Desktop apps, Devin, and running agents in parallel 43:30 Remote environments and working in monorepos 44:34 Skills, plugins, and hooks 47:57 Parallel worktrees and staying focused 50:42 Agent teams vs subagents 53:06 The Finch mission and wrap up

    Siddharth Nanda - Microsoft Engineer Reveals how Engineering Will Never be the Same Again
  8. 30 jun

    Voice AI - The Next Frontier | Decagon, Retell, Vapi, Smallest AI, Daily

    Voice agents are one of the hottest use cases in enterprise right now, but also one of the hardest to actually take live. Getting latency low enough to feel human without dumbing down the responses, making reliable tool calls to a CRM without dropping the customer mid-call, building fallback models for when Anthropic or OpenAI are running hot. None of it is as simple as the demos make it look. Last week I hosted a fireside chat with five eng leaders who deal with this stuff every day: Basia Sudol (Head of Enterprise Solutions, Decagon), Varun Singh (CPTO, Daily), Steven Diaz (FDE Manager, Vapi), Tyler D'Silva (Founding FDE, Retell AI), and Sudarshan Kamath (Founder, Smallest AI). We get into why nobody serious is shipping real-time voice-to-voice yet, why LLMs forget the middle of your prompt (and what that does to your architecture), why a giant prompt quietly destroys your unit economics, and why voice agent costs are now being compared directly against human labor. Plus the stuff nobody warns you about: turn-taking, HIPAA constraints, why outbound is easier than inbound, why getting an exec to actually like the voice can be harder than any model problem, and more! Chapters below: 00:00 Intro 00:26 Meet the panel 01:23 Daily, WebRTC, and 20 years of building voice 03:49 How Smallest AI made real-time TTS work 05:23 Why Decagon moved into voice 08:16 How Vapi and Retell think about the stack 11:14 Forward deployed vs solutions engineering 16:48 Voice agent architecture, explained simply 22:11 Cascade vs speech-to-speech: the real tradeoff 28:11 Hybrid pipelines and mixing models 32:16 Accents, multilingual, and getting Singlish right 35:32 Prompts vs workflows, and the latency fight 44:26 How you actually evaluate a voice agent 49:04 Simulation-based evals 49:49 What production metrics really look like 51:53 Building a QA framework that scales 54:53 Evaluating speech-to-speech 58:10 Open source benchmarks 59:57 Why picking a voice is so subjective 01:02:03 Personalization and custom voices 01:03:20 Voice quality is solved, GPU efficiency is the new war 01:06:23 Why outbound calls work better than you'd think 01:10:38 Deploying in regulated industries (HIPAA, retention, audits) 01:12:43 Turn-taking, the hardest unsolved problem in voice 01:18:53 Where voice agents go in the next year 01:27:55 Audience Q&A: inside Smallest's Hydra model 01:32:15 The deployment problems nobody has solved yet 01:37:46 Closing thoughts and thanks

    Voice AI - The Next Frontier | Decagon, Retell, Vapi, Smallest AI, Daily

Acerca de

Discover how leading enterprises and professionals turn AI into real products. Hear candid conversations with executives and builders who deploy AI at scale and learn what works (and what doesn't).

También te podría interesar