The Experimentation Edge

Growthbook

How do product teams decide what to build and what not to? The Experimentation Edge is the podcast where product, growth, and engineering leaders share how A/B testing, feature flags, and experimentation drive real business outcomes — backed by named companies and real numbers. From DoorDash's 12,000 A/B tests a year to Atlassian's experimentation-led product win to UPS's $500M experimentation team, each episode goes deep with operators running experimentation programs at scale. Hosted by Ashley Stirrup, CMO at GrowthBook and a 25-year executive in data and experimentation. For product managers, engineers, data scientists, and growth leaders at B2B tech companies who care about experimentation culture, statistical rigor, and shipping with confidence. No marketing speak. Just operators explaining what they shipped, what moved the needle, and how experimentation reshaped their teams. Topics: A/B testing, experimentation, growth experimentation, product experimentation, tech experimentation, feature flags, experimentation culture, statistical significance, marketplace experimentation, conversion rate optimization, experimentation at scale.

  1. 19h ago

    Realtor.com on using your AI as a junior data scientist

    Summary What do you do when your biggest experiment win turns out to be a loss? Whitney Perez, Director of Product Management at Realtor.com, joins host Ashley Stirrup to share the checkout bundling test that posted a 300% attach rate and still lost revenue, the 30/30/30 rule she uses to set expectations for a new experimentation team, and how AI is turning an English major into an aspirational data scientist. This episode is for product managers, engineers, and data scientists building experimentation programs from the ground up. Chapters 00:00 Cold open and welcome 01:10 From growth hacker to Realtor.com 02:50 Three foundations for a new experimentation team 04:20 The 30/30/30 rule 05:05 The 300% bundling win that lost revenue 07:10 You don't need a stats degree to experiment 09:20 Cascading North Star metrics 12:50 Do the homework before the experiment 13:50 The wishlist: instrumentation, embedded knowledge, culture 15:50 AI as an aspirational data scientist 18:45 Keeping a human in the loop Takeaways -A winning decision metric is not enough. Realtor.com's bundling test hit a 300% attach rate, but funnel fallout from the extra step made it a net revenue loser. Set secondary metrics and their thresholds before launch. -Expect the 30/30/30 rule: roughly a third of tests win, a third are inconclusive, and a third lose. The math is the math, and the losers carry most of the learning. -Start a new team on foundations: what a clean test and an A/A test look like, which surfaces should not be tested, and a peer review program that lets people graduate to more complex experiments. -You don't need a stats background to run good experiments. Teach the simplest definition of a good test, then let people learn by doing. -AI can make anyone an aspirational data scientist for analyzing results and spotting opportunities, but it can be confidently wrong. Keep a human in the loop and sanity check output the way you'd peek at a freshly launched test. Connect with the Guest LinkedIn: https://www.linkedin.com/in/whitneykperez/ Website: https://www.realtor.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    Realtor.com on using your AI as a junior data scientist
  2. 1d ago

    How Zalando connects every experiment to its North Star

    Summary How do you keep 1,000 experiments a year pointed at one North Star? In this episode of The Experimentation Edge, host Ashley Stirrup talks with Mi Tian, Head of Applied Science at Zalando, about running experimentation inside a central economics org that reports to the CFO. Mi shares how Zalando balances safe confirmatory tests with game-changing bets, how the team measured the discovery feeds homepage launch when success had no established metric, and how a KPI tree cascades the company North Star down to the controllable inputs teams ship every day. She also looks ahead to LLM-based agents as a simulation layer for screening hypotheses. A practical conversation for anyone building an experimentation program that wants both rigor and ambition. Chapters 00:45 About Zalando and its marketplace model 02:00 Mi's path from engineering to experimentation 03:10 Economists and data scientists in one decision-making org 04:15 Running over 1,000 experiments a year 05:55 What makes an experiment high risk 07:10 Sharing learnings through standardization and champions 09:10 The discovery feeds launch and its measurement plan 11:15 Balancing the experimentation portfolio 12:35 Growing a KPI tree from the North Star 16:05 LLM agents and the future of experimentation at Zalando 18:35 New missions for a longstanding business Takeaways -Treat experimentation as a portfolio: balance confirmatory tests that protect the business with game-changing bets that can win big. -Assess risk tiers when building the roadmap so measurement rigor scales with the stakes instead of slowing every decision down. -When a launch is too new to have a success metric, pair short term A/B tests with long term holdouts from day one. -Connect every experiment to the company North Star by cascading it down to sensitive proxy metrics and controllable inputs. -Use LLM based agents as a cheap simulation layer to screen hypotheses, not as a replacement for real A/B tests. Connect with the Guest LinkedIn: https://www.linkedin.com/in/mi-tian-941b4767/ Website: https://www.zalando.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    How Zalando connects every experiment to its North Star
  3. 2d ago

    Learneo on testing the opposite of every hypothesis

    SummaryRich Liebling, senior director of engineering at Learneo, joins host Ashley Stirrup to explain the practice that came out of growing Shop It To Me from 50,000 subscribers to one million in nine months: test your hypothesis, and test its opposite. Rich covers the page where cutting text lost and adding text won, why the inverse wins more often than teams expect, and why small focused tests are the only ones where "the opposite" means anything. He also walks through translating a million subscriber goal into a target of 300 A/B tests, making a new engineer's second merge request their own experiment, and the different constraints at Course Hero, where competing team metrics were resolved with an early lifetime value model and three-day SQL analyses quietly capped testing velocity. This episode is for product managers, engineers, data scientists, and growth leaders building or scaling an experimentation program. Chapters 00:00 Cold open and introduction 01:50 Shop It To Me and the first engineering hire 02:55 300 A/B tests as the path to one million subscribers 04:00 Testing as a core value and the second merge request 07:20 Learneo, Course Hero, and two ways to get access 08:50 Modeling lifetime value to stop teams competing 11:10 Testing the opposite of the hypothesis 14:35 A portfolio of small, medium, and large tests 17:15 Multi armed bandits and seasonal traffic 20:35 The flywheel that keeps copycats behind Takeaways -Test the opposite of every hypothesis. At Shop It To Me the inverse won surprisingly often, and even when it lost it proved the variable mattered. -Keep tests small and focused. Redesign a whole page and lose, and you learn that version failed but not what to change next. -Translate a growth goal into an execution count. One million subscribers is not actionable. 300 A/B tests by year end is, and everyone can influence it. -Make experimentation part of hiring and onboarding. Every new engineer's second merge request was their own test idea. -Analysis friction sets the ceiling on testing velocity. Three days of ad hoc SQL per test quietly discourages teams from running more. Connect with the GuestLinkedIn: https://www.linkedin.com/in/richliebling/Website: https://www.learneo.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    Learneo on testing the opposite of every hypothesis
  4. Sep 3

    The four questions Early Warning asks before any A/B test

    Summary What separates a valid A/B test from an expensive guess? Priya Singhee, VP of Enterprise Analytics & Data Science at Early Warning — the bank-owned consortium that fights payment fraud and operates Zelle, which processed a trillion dollars last year — joins host Ashley Stirrup to share the experimentation playbook she built leading storefront analytics at Wayfair. She walks through the four questions to ask before launching any A/B test, why 85 to 90% of tests are supposed to fail, how pre-registration and kill criteria stop p-hacking before it starts, the pitfalls that fake wins (novelty effects, hidden heterogeneity, multiple comparisons), and how to roll out winners with gradual ramps and long-running holdouts. A practical episode for product managers, engineers, data scientists, and growth leaders building rigorous experimentation programs. Chapters 00:45 Meet Early Warning: fraud detection, Zelle, and a trillion dollars in payments 02:00 Wayfair and optimizing every step of the storefront funnel 03:05 The four questions to ask before any A/B test 05:15 Test setup best practices: hypotheses, guardrails, power, and pre-registration 07:40 Why 85 to 90% of tests fail and why that's a learning agenda 09:25 Novelty effects, hidden heterogeneity, and the multiple comparisons problem 12:55 Pre-registration, kill criteria, and stopping p-hacking 15:10 Rolling out winners: gradual ramps and long-running holdouts 17:15 Causal inference when you can't A/B test 20:05 The case for more A/B testing, not less Takeaways Run the four-question framework before any test: clean randomization, a plausible effect size for your traffic, a reversible and cheap change, and a falsifiable hypothesis. Treat A/B testing as a learning agenda: 85 to 90% of tests are supposed to fail, and a suspiciously high win rate is a red flag, not a trophy. Pre-register the full analysis plan, including hypothesis, mechanism, primary metric, exact statistical test, and subgroups, so p-hacking can't creep in when a test goes sideways. Define kill criteria and success, failure, and guardrail-dip actions before launch, with leadership sign-off, so nobody chases a loss into a fake win. Log every test and its learnings where the whole organization can see them; that reinforcement loop is what separates world-class experimentation programs. Connect with the Guest LinkedIn: https://www.linkedin.com/in/priya-singhee/ Website: https://www.earlywarning.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    The four questions Early Warning asks before any A/B test
  5. Sep 1

    How Supercell A/B tests 300 million players without breaking trust

    SummaryOn this episode of The Experimentation Edge, Ashley Stirrup talks with Shan Huang, data scientist on the central experimentation team at Supercell, the Helsinki mobile game company behind Clash of Clans, Clash Royale, Brawl Stars, Hay Day, and Boom Beach. Shan explains how a famously decentralized, creative-first company with 300 million monthly active users runs fewer than 100 A/B tests a quarter and why that restraint is deliberate, how Supercell announces experiments to players in advance and promises make-up events to keep testing fair, and how importable AI skills now let anyone at the company analyze their own experiments, making quality consistency the next big challenge. It's for product managers, data scientists, and growth leaders balancing creative conviction with experimental rigor. Chapters 00:00 Intro 01:10 About Supercell and 300 million players 02:30 The central team and a decentralized culture 04:25 Fewer than 100 tests a quarter 05:40 Sharing learnings across independent game teams 07:35 Retention as the North Star 08:35 Onboarding experiments with gems and tutorials 12:45 Telling players about A/B tests 15:15 Hypotheses and proxy metrics 20:45 AI and the future of experiment analysis Takeaways -Supercell runs fewer than 100 A/B tests a quarter for 300 million monthly players, because the goal is to become more hypothesis driven while staying creative, not to maximize volume. -In a decentralized company, a central experimentation team earns its impact by providing the platform, partnering on rigor, and making sure learnings travel across independent game teams. -Announce experiments to users before they run; Supercell's community managers tell players what is being tested and why, which turns a skeptical community into a research partner. -Promise fairness, not just transparency; players who get the worse variant always receive a make-up event later, because game players come to have fun, not to be disadvantaged. -AI-powered self-serve analysis means everyone can now run and analyze experiments, so the next challenge is making the quality of AI analysis consistent across the whole company. Connect with the GuestLinkedIn: https://www.linkedin.com/in/cnshanhuang/Website: https://supercell.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    How Supercell A/B tests 300 million players without breaking trust
  6. Aug 27

    How Clover experiments when billions of dollars flow through daily

    Summary How do you run an experimentation program when classic A/B testing is off the table? Ben Schein, Director of Product Management at Clover, joins host Ashley Stirrup to explain how a platform serving 300,000+ merchants and processing billions of dollars daily proves every feature through pilots and ground-level testing before rollout, why uncertainty and downside — not feature visibility — decide testing depth, and what his years leading product at Shake Shack taught him about turning a checkout funnel into a brand channel. This episode is for product managers, engineers, and data scientists building experimentation programs where the stakes are real. Chapters 00:00 Cold open and welcome 01:30 The Clover business model and its scale 03:45 Ben's role and the metrics that matter 07:45 Deciding what gets tested: uncertainty and downside 09:05 Why Clover can't test in production 12:45 Testing the Shake Shack checkout experience 19:35 Advice for PMs new to experimentation 23:45 Context over personalization 25:45 The future of experimentation and the human element Takeaways -At Clover's scale, testing happens before rollout: pilots and detailed go-to-market plans replace in-production A/B tests, because a merchant's work tool can never change overnight without warning. -Uncertainty and downside set the testing depth. High-risk changes like payment authorization flows earn deep, ground-level experimentation, while table stakes features like Apple Pay earn a monitored rollout. -Testing is the evidence that justifies rollout investment: if the data doesn't show a feature will succeed, the go-to-market dollars never get spent. -A checkout funnel can carry the brand. At Shake Shack, Ben's team tested prep-time expectations, fixed wrong-location orders, and used loading screens to deliver hospitality digitally. -Structure experiments for durable business value, not pass-fail verdicts: pair headline metrics with counter metrics and anchor on ground-level measurements like items per check that resist marketing noise. Connect with the Guest LinkedIn: https://www.linkedin.com/in/benschein/ Website: https://www.clover.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    How Clover experiments when billions of dollars flow through daily
  7. Aug 26

    Why JobLeads says one test won't move you, but 100 will

    SummaryWhat happens when the most logical feature you've ever built has zero impact? In this episode of The Experimentation Edge, host Ashley Stirrup, CMO of GrowthBook, sits down with Edd Saunders, product experimentation manager at JobLeads, to unpack the pizza personalization experiment that cut an ordering flow from 22 clicks to 5 and changed nothing. Edd shares the problem mapping framework he uses to move new experimenters from solution space to problem space thinking, how JobLeads grew from 0.3 to 2.8 experiments per month, and why democratizing experimentation across a whole company comes down to habit change rather than education. A practical conversation for product managers, data scientists, engineers, and growth leaders building experimentation cultures. Chapters 00:00 Introduction 01:40 Meet Edd Saunders and JobLeads 03:12 How JobLeads uses AI for prototyping 04:05 Building experimentation operations and a knowledge base 05:35 Learning over winning and compounding growth 08:10 The pizza personalization experiment 12:50 Moving from solution space to problem space 17:05 Problem mapping on a 2x2 matrix 22:20 Velocity and democratizing experimentation 24:20 AI automation for the unsexy work Takeaways -A one click reorder feature that cut a pizza ordering flow from 22 inputs to 5 had zero impact on purchases, proving that removing friction can also remove the customer's sense of control. -Exploration is part of the customer's delight; returning customers wanted to browse the menu even though they ordered the same thing every week. -Moving new experimenters from solution space to problem space thinking raises win rates and produces learnings the whole organization can use. -Problem mapping on a 2x2 matrix of evidence versus impact turns customer research into a prioritized experiment roadmap, and one validated problem can spring a whole tree of testable ideas. -Scaling experimentation from 0.3 to 2.8 tests per month is less about education and more about habit change, shared learnings, and giving non specialists the tools to launch their own experiments. Connect with the GuestLinkedIn: https://www.linkedin.com/in/eddsaunders/Website: https://www.jobleads.comRead Blog: https://www.growthbook.io/podcast/episode/1-34 SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    Why JobLeads says one test won't move you, but 100 will
  8. Aug 20

    Why The Aspen Group targets a 25% win rate

    SummaryWhat does a winning A/B test mean when your website serves 1,100 dentist owned offices? In this episode of The Experimentation Edge, host Ashley Stirrup talks with Arie Polycarpou, Senior Manager of Test and Learn at Aspen Dental, about building experimentation programs at Kohl's, Marriott, Total Wine, and now the largest company in The Aspen Group's healthcare retail portfolio. Arie explains why appointment bookings are the North Star but never the whole story, why he deliberately targets a 25% win rate and expects it to fall as the program matures, and why he applies a haircut to every stacked win before it reaches leadership. A practical conversation for product managers, engineers, data scientists, and growth leaders building test and learn programs at multi-location businesses. Chapters 00:00 introduction and Arie's path into experimentation 02:05 building programs at Kohl's, Marriott, and Total Wine 03:15 healthcare retail and 1,100 dentist owned offices 04:30 the test and learn team and 100 tests a year 07:15 building a culture of shared wins and explained losses 09:50 what a losing navigation redesign revealed 12:30 testing your way into big redesigns 13:55 the case for a 25% win rate 15:50 haircuts, holdouts, and honest math on stacked wins 17:55 metrics beyond conversion and where AI fits next Takeaways -Aspen Dental runs experimentation as healthcare retail: with 1,100 dentist owned offices, the office, not just the website visitor, is the real unit of analysis. -A mature program should target a true win rate around 25%; a 40% win rate usually signals a young program still picking off low-hanging fruit. -Apply a haircut to stacked wins: lifts depreciate as customers acclimate, and two 5% wins never add up to 10%. -Losing tests are valuable when they're designed to isolate why: test your way into big redesigns instead of shipping them whole. -Experimentation culture grows from sharing wins and explaining losses: keep the statistical rigor in the back end and the communication simple. Connect with the GuestLinkedIn: https://www.linkedin.com/in/ariepolycarpou/Website: https://www.aspendental.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    Why The Aspen Group targets a 25% win rate

About

How do product teams decide what to build and what not to? The Experimentation Edge is the podcast where product, growth, and engineering leaders share how A/B testing, feature flags, and experimentation drive real business outcomes — backed by named companies and real numbers. From DoorDash's 12,000 A/B tests a year to Atlassian's experimentation-led product win to UPS's $500M experimentation team, each episode goes deep with operators running experimentation programs at scale. Hosted by Ashley Stirrup, CMO at GrowthBook and a 25-year executive in data and experimentation. For product managers, engineers, data scientists, and growth leaders at B2B tech companies who care about experimentation culture, statistical rigor, and shipping with confidence. No marketing speak. Just operators explaining what they shipped, what moved the needle, and how experimentation reshaped their teams. Topics: A/B testing, experimentation, growth experimentation, product experimentation, tech experimentation, feature flags, experimentation culture, statistical significance, marketplace experimentation, conversion rate optimization, experimentation at scale.

You Might Also Like