The Experimentation Edge

Growthbook

How do product teams decide what to build and what not to? The Experimentation Edge is the podcast where product, growth, and engineering leaders share how A/B testing, feature flags, and experimentation drive real business outcomes — backed by named companies and real numbers. From DoorDash's 12,000 A/B tests a year to Atlassian's experimentation-led product win to UPS's $500M experimentation team, each episode goes deep with operators running experimentation programs at scale. Hosted by Ashley Stirrup, CMO at GrowthBook and a 25-year executive in data and experimentation. For product managers, engineers, data scientists, and growth leaders at B2B tech companies who care about experimentation culture, statistical rigor, and shipping with confidence. No marketing speak. Just operators explaining what they shipped, what moved the needle, and how experimentation reshaped their teams. Topics: A/B testing, experimentation, growth experimentation, product experimentation, tech experimentation, feature flags, experimentation culture, statistical significance, marketplace experimentation, conversion rate optimization, experimentation at scale.

  1. 5d ago

    Why US Bank considers missing even 1% of customers unacceptable

    SummaryHow does a major bank scale experimentation when even one percent of customers missing an experience is unacceptable? Vijay Lal, Lead Product Manager for Experimentation at US Bank, joins host Ashley Stirrup, CMO at GrowthBook, to share how his team made their experimentation platform self serve for non technical marketers, how a login widget experiment led to a two second fallback that accounted for every customer, and why metrics should be driven by hypotheses instead of handed down by leadership. They also dig into where AI genuinely saves time in experiment analysis, why a human in the loop is non negotiable, and what real time personalization means for the future of testing. This episode is for product managers, data scientists, and experimentation leaders, especially those working in regulated industries. Chapters 00:00 Cold open and welcome 00:40 Vijay's path from Comcast to financial services 03:16 Making the experimentation platform self serve 05:08 The login widget experiment and the two second fallback 08:59 Documenting learnings from every experiment 10:38 AI in experimentation and the human in the loop 12:39 Advice for new product managers 15:24 Hypothesis driven metrics 18:25 Real time personalization and agentic AI 20:11 Democratizing experimentation with responsibility Takeaways -Self serve experimentation lets a small central team support a huge testing volume, but it only works with continuous training and guardrail metrics attached. -In a regulated industry, every customer must be accounted for. Even one to two percent of users missing an experience is unacceptable. -A simple fallback, like a two second load rule, can save an ambitious experiment without sacrificing coverage or security. -Metrics should be driven by the experiment's hypothesis, not chosen by leadership in a silo. Pair a primary KPI with secondary KPIs for return behavior. -AI saves real time in experiment analysis, but a human in the loop must validate anything AI produces before it goes live. Connect with the GuestLinkedIn: https://www.linkedin.com/in/vijay-lal/ Website: https://www.usbank.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    Why US Bank considers missing even 1% of customers unacceptable
  2. Aug 5

    Why Farfetch manages by learning rate, not win rate

    Summary Luis Trindade, Principal Product Manager of Experimentation at Farfetch, joins host Ashley Stirrup to explain how one of the world's largest luxury marketplaces built its own experimentation platform and the culture around it. Luis covers the move from a hybrid setup with an external testing vendor to Fabs 2.0, the in house system where a feature toggle is the single entry point for every experiment, why Farfetch manages by learning rate instead of win rate, and the two year Inspire experiment that replaced the world's leading recommendation engine vendor. He also shares how a deliberately shrinking center of excellence supports hundreds of experiments a month through clinics, shared templates, and open learning sessions. This episode is for product managers, engineers, data scientists, and growth leaders building or scaling an experimentation program. Chapters 00:00 Cold open and introduction 01:45 Inside Farfetch, the global marketplace for luxury fashion 08:00 From startup validation to an experimentation mindset 09:45 A center of excellence that enables instead of executes 12:45 Fabs, build versus buy, and dropping the external vendor 16:45 One feature toggle as the entry point for every experiment 20:45 Learning rate over win rate 23:15 The two year experiment that replaced the recommendation vendor 29:45 Onboarding new product managers into experimentation 33:15 AI, corporate knowledge, and what comes next for experimentation Takeaways -Manage by learning rate, not win rate. The only failed test is one that was badly designed, with wrong metrics or sampling biases. Every other test produces a learning. -Route every experiment through a single entry point. Farfetch's feature toggling system connects segmentation, user systems, CMS, and messaging so every team tests in the same language. -External JavaScript injection tools carry hidden costs: broken pages, inconsistent results, and rework to reclaim your own data for deep dives. -Strategic bets deserve a longer clock than fail fast allows. Farfetch iterated on its Inspire engine for two years before it beat and replaced the market leader. -A center of excellence should enable, not execute. Farfetch's central team shrank while experiment volume grew because its job is ceremonies, templates, and coaching. Connect with the Guest LinkedIn: https://www.linkedin.com/in/ltrindade/ Website: https://www.farfetch.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    Why Farfetch manages by learning rate, not win rate
  3. Jul 23

    How Cogniteer Built an Experimentation Engine From Scratch

    Summary In this episode of The Experimentation Edge, host Ashley Stirrup talks with Fabian Hans, founder and behavioral psychologist at Cogniteer, a consultancy that helps enterprises build in-house experimentation programs and raise both test velocity and win rate. Drawing on fifteen years in conversion rate optimization, Fabian explains why mass-producing the same A/B tests across clients quietly kills learning, why most ecommerce drop-offs are structural rather than your fault, and how matching the interface to how people actually buy, new versus returning, B2C versus B2B, can move conversion far more than another button. It is a practical, psychology-grounded conversation for product managers, engineers, data scientists, and growth leaders who want their experimentation programs to compound understanding, not just volume. Chapters 00:00 Introduction 01:25 From agency mass production to in house deep dives 04:05 Why some products resist selling online 06:35 The drop offs every ecommerce shop shares 07:45 The 50% win rate test Cogniteer reused 09:25 Why alignment beats developer resources 12:45 Two teams, two goals, one broken checkout 17:05 Selling water dispensers without a product catalog 22:15 Designing every experiment to lose 26:15 Personalizing buyers and where AI takes experimentation Takeaways -Deep dives beat mass produced tests, because understanding one business's users uncovers bigger levers than reusing the same test across many clients. -Many ecommerce drop offs are structural, since the basket and product page leak in roughly 80% of shops because it is ecommerce, not because of your product. -Product to channel fit decides what sells online, so books and fashion judge well on a screen while perfume and washing machines need cues the interface cannot fully provide. -The real bottleneck is alignment, not developer resources, so agree on the problem and its hierarchy before anyone builds a variation. -Match the interface to how people actually buy, because new buyers need information, returning buyers want speed, and B2B buyers often want a solution and an offer instead of a product catalog. Connect with the Guest LinkedIn: https://www.linkedin.com/in/fabianhans-cogniteer/ Website: https://www.cogniteer.de/ SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    How Cogniteer Built an Experimentation Engine From Scratch
  4. Jul 21

    How Fin A/B Tests Millions of Samples in Days

    Summary On this episode of The Experimentation Edge, host Ashley Stirrup talks with Pedro Tabacof, Principal Machine Learning Scientist at Fin (formerly Intercom), about how one of the most advanced AI customer support agents in the world is built on relentless experimentation. Pedro explains why unit tests don't work on non-deterministic AI, how Fin runs up to two dozen concurrent A/B tests pulling millions of samples in days, and shares two counterintuitive experiments: one where slowing the agent down improved every metric, and one where adding more context made Fin more helpful and more prone to fake promises until a targeted prompt fix kept the upside without the hallucinations. It's a candid look for product managers, engineers, and data scientists at how a $100M ARR AI product actually ships improvements. Chapters 00:00 Welcome and what Fin actually does 02:00 How Fin became Anthropic's first line of support 02:30 Why Fin sells resolutions not deflections 06:00 Owning the stack with custom models 10:40 Pedro's path from fuzzy logic to AI 12:55 Why A/B testing is the only gold standard for AI 15:50 Do no harm testing on every change 18:00 The latency experiment that shocked the team 27:30 When more context made Fin hallucinate 30:15 Win rates and the future of AI driven experimentation Takeaways -Faster is not always better. Fin increased latency artificially and positive feedback went up, likely because a small delay makes an AI feel like it is doing real work. -You cannot unit test a non-deterministic AI. A/B testing at scale, millions of samples in days, is the only reliable way to know a change actually helped. -Adding more conversation history made Fin more helpful and more prone to fake promises, until a targeted prompt fix removed the hallucinations and kept most of the gain. -A losing experiment is often a winner with one broken part. Diagnose which element hurts the experience, fix only that, and rerun. -Fin A/B tests everything, even one-character prompt changes and many bug fixes, and treats a 20 to 30 percent win rate as a healthy sign of a real experimentation program. Connect with the Guest LinkedIn: https://www.linkedin.com/in/tabacof/ Website: https://fin.ai SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    How Fin A/B Tests Millions of Samples in Days
  5. Jul 14

    How Kargo turns losing experiments into competitive edges

    Summary In this episode of The Experimentation Edge, host Ashley Stirrup, CMO of GrowthBook, sits down with James Falzone, Director of Product Management at Kargo, to unpack how a high scale ad tech marketplace turns failure into its biggest advantage. James explains how Kargo connects advertisers to publishers through real time auctions that resolve in milliseconds across up to 10 billion ad requests a day, why experimentation is embedded in the company's culture rather than siloed in a team, and what happened when a winning click optimization model failed completely after being copied to a new customer type. The conversation is built for product managers, data scientists, engineers, and growth leaders who want a practical, honest view of running experiments at scale, learning from losses, and keeping AI grounded in solid infrastructure. Chapters 00:00 Welcome and introducing James Falzone 01:45 What Kargo does and how real time ad auctions work 04:45 Why experimentation is embedded in Kargo's culture 07:45 The three things every marketplace has to deliver 10:15 The experiment that failed: click optimization on third party demand 12:15 A bad result versus a bad experiment 13:45 Why different customer types need different signals 15:30 Putting "where did you fail?" on every retro 18:45 How experimentation evolves with AI 21:15 Better not bigger: the closing takeaway Takeaways -A bad result is not a bad experiment. If you're not failing, you're probably not trying anything new. -The same metrics and signals don't apply to every customer type. Bad results often come from a lack of context, not bad tech. -Metrics and signals you test against should always be business driven, not ported from the last thing that worked. -Put failure on the agenda. A biweekly "where did you fail?" retro turns one person's dead end into the whole team's shortcut. -AI's biggest unlock is access. More people can run experiments, but it has to be built on solid ML and infrastructure. Better, not bigger. Connect with the Guest LinkedIn: https://www.linkedin.com/in/jamesafalzone/ Website: https://kargo.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    How Kargo turns losing experiments into competitive edges
  6. Jul 9

    The 'wine effect' and other surprises that reshaped how Box runs e-commerce experiments

    Summary In this episode of The Experimentation Edge, host Ashley Stirrup talks with Danielle Olean, Director of E-commerce at Box, about what it really takes to build a culture of experimentation inside a B2B company. Drawing on 15 years across B2C and B2B at Wayfair, Drizly, Zoom, and now Box, Danielle explains why experimentation belongs to every product team and not just e-commerce, walks through a pricing page saga of one win and two losses that exposed the limits of simplification, and shares the "wine effect" test that won for a reason no one predicted. It's a practical, story rich conversation for product managers, growth leaders, and anyone trying to make better decisions with data. Chapters 00:45 Meet Danielle Olean and Box's reinvention 02:45 Owning the entire customer life cycle 04:45 Why experimentation matters even without a checkout 07:45 The feature that's used but hidden 11:45 Proving ROI with a scrappy manual test 12:45 Building a culture that shares wins and losses 16:45 The pyramid strategy for prioritizing tests 18:45 The simplification tightrope on the pricing page 24:45 When a test wins for the wrong reason 27:45 Where experimentation at Box goes next Takeaways - Experimentation isn't only for e-commerce. Any product with a funnel, even an AI chatbot, can be measured and improved through testing. - Simplification has a limit. Removing too much can strip away the cues and context buyers actually need to decide. - Share losses as openly as wins. Wins build credibility, and losses build the psychological safety a testing culture runs on. - Prioritize like a pyramid. Fix the widest-impact experiences first, then optimize down into smaller cohorts. - Surprising results are the point. A test can win for a reason you never hypothesized, like the "wine effect," and that's where the real learning lives. Connect with the Guest LinkedIn: https://www.linkedin.com/in/dolean1/ Website: https://www.box.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    The 'wine effect' and other surprises that reshaped how Box runs e-commerce experiments
  7. Jul 7

    Dilligent explains why moving on from an experiment might cost you

    Summary Dan Layfield, Director of Product Management at Diligent, joins host Ashley Stirrup on The Experimentation Edge to trace what fifteen years of A/B testing across Codecademy, Uber Eats, and the Fortune 1000 boardroom actually taught him. He breaks down the Codecademy trial-model rebuild that took four months and several rounds to deliver a 35% conversion lift, why moving on from a losing experiment too early is one of a PM's costliest mistakes, how to escape the B2B feature factory with metrics that genuinely ladder up, why retention should ride a product's natural use case instead of fighting it, and where AI is already replacing weeks of research and analysis. It's a practitioner's guide for product managers, growth leaders, data scientists, and engineers bringing experimentation rigor to both B2C and B2B. Chapters 00:45 Meet Dan Layfield and Diligent 01:45 Two worlds of experimentation, Codecademy and Uber 03:45 The trial model that lifted conversion 35% 06:20 What to do with a losing experiment 08:50 Two flavors of experimentation 09:45 Reading forty metrics at Uber Eats 13:10 Escaping the B2B feature factory 16:45 Anchoring the North Star to real usage 19:15 Where AI fits in research and analysis Takeaways A losing experiment is often inconclusive, not negative; treat it as a map of the funnel rather than a verdict, and know when a big problem is worth another round.Persistence paid off at Codecademy: four months and three to four rounds of trial-model testing produced a 35% conversion increase.Separate your two experimentation modes; high-volume CRO chases many small wins, while big, uncertain bets are worth taking multiple shots to de-risk.Most B2B product teams are feature factories; the fix is a top-down OKR system, and planning usually breaks in the connections between layers, not inside them.Anchor retention and engagement to the product's natural use case, and use AI to synthesize research and simple A/B analysis in hours instead of weeks.Connect with the Guest LinkedIn: https://www.linkedin.com/in/layfield/ Website: https://www.diligent.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    Dilligent explains why moving on from an experiment might cost you
  8. Jul 2

    The metric Stitch Fix says every experimenter should chase

    Summary In this episode of The Experimentation Edge, GrowthBook CMO Ashley Stirrup sits down with Nick Beyler, data science manager at Stitch Fix, where he leads the decision and insights team and owns the company's internal experimentation platform. Nick shares why the metric he most wants is the one he can't measure yet, a North Star that predicts a client's long-term value from their earliest behaviors, and why the most impactful experiment learnings tend to come from adoption friction rather than product bugs. He makes the case that if you're only testing winners you're not taking enough risks, explains how guardrails make that risk safe, and looks ahead to a new in-house platform and the promise of agentic AI. It's a practical, statistician's-eye view of experimentation for product managers, data scientists, and engineers building serious testing programs. Chapters 00:00 Cold open and welcome to the show 01:45 What Stitch Fix actually does 04:15 Balancing AI with the human stylist 05:15 From public policy to the A/B testing adrenaline rush 07:15 Inside the weekly experimentation review group 08:45 The AI style assistant and listening to qualitative feedback 10:45 Why adoption friction beats product bugs 13:45 Testing for losers and building guardrails 15:45 Keep rate, successful fixes, and the holy grail metric 18:15 The new platform and the promise of agentic AI Takeaways The most impactful experiment learnings usually come from adoption friction, not product bugs. By the time a big feature reaches A/B testing, it's often already a winner, so the open question is how and where to introduce it.A losing test is a finding, not a failure. If every experiment wins, you're not taking enough risk to learn anything new.Guardrails and stopping criteria are what make risk-taking safe, especially when the experience is as personal as shopping.The most valuable North Star metric is the one you can't measure yet, long-term client value, and causal-inference modeling helps predict it from short-term behavior.Quantitative results are only half the story. Direct, qualitative client feedback inside an experiment often reshapes the rollout more than the numbers do.Connect with the Guest LinkedIn: https://www.linkedin.com/in/nick-beyler-381864119/ Website: https://www.stitchfix.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    The metric Stitch Fix says every experimenter should chase

About

How do product teams decide what to build and what not to? The Experimentation Edge is the podcast where product, growth, and engineering leaders share how A/B testing, feature flags, and experimentation drive real business outcomes — backed by named companies and real numbers. From DoorDash's 12,000 A/B tests a year to Atlassian's experimentation-led product win to UPS's $500M experimentation team, each episode goes deep with operators running experimentation programs at scale. Hosted by Ashley Stirrup, CMO at GrowthBook and a 25-year executive in data and experimentation. For product managers, engineers, data scientists, and growth leaders at B2B tech companies who care about experimentation culture, statistical rigor, and shipping with confidence. No marketing speak. Just operators explaining what they shipped, what moved the needle, and how experimentation reshaped their teams. Topics: A/B testing, experimentation, growth experimentation, product experimentation, tech experimentation, feature flags, experimentation culture, statistical significance, marketplace experimentation, conversion rate optimization, experimentation at scale.

You Might Also Like