The Experimentation Edge

Growthbook

How do product teams decide what to build and what not to? The Experimentation Edge is the podcast where product, growth, and engineering leaders share how A/B testing, feature flags, and experimentation drive real business outcomes — backed by named companies and real numbers. From DoorDash's 12,000 A/B tests a year to Atlassian's experimentation-led product win to UPS's $500M experimentation team, each episode goes deep with operators running experimentation programs at scale. Hosted by Ashley Stirrup, CMO at GrowthBook and a 25-year executive in data and experimentation. For product managers, engineers, data scientists, and growth leaders at B2B tech companies who care about experimentation culture, statistical rigor, and shipping with confidence. No marketing speak. Just operators explaining what they shipped, what moved the needle, and how experimentation reshaped their teams. Topics: A/B testing, experimentation, growth experimentation, product experimentation, tech experimentation, feature flags, experimentation culture, statistical significance, marketplace experimentation, conversion rate optimization, experimentation at scale.

  1. 18h ago

    How Zalando connects every experiment to its North Star

    Summary How do you keep 1,000 experiments a year pointed at one North Star? In this episode of The Experimentation Edge, host Ashley Stirrup talks with Mi Tian, Head of Applied Science at Zalando, about running experimentation inside a central economics org that reports to the CFO. Mi shares how Zalando balances safe confirmatory tests with game-changing bets, how the team measured the discovery feeds homepage launch when success had no established metric, and how a KPI tree cascades the company North Star down to the controllable inputs teams ship every day. She also looks ahead to LLM-based agents as a simulation layer for screening hypotheses. A practical conversation for anyone building an experimentation program that wants both rigor and ambition. Chapters 00:45 About Zalando and its marketplace model 02:00 Mi's path from engineering to experimentation 03:10 Economists and data scientists in one decision-making org 04:15 Running over 1,000 experiments a year 05:55 What makes an experiment high risk 07:10 Sharing learnings through standardization and champions 09:10 The discovery feeds launch and its measurement plan 11:15 Balancing the experimentation portfolio 12:35 Growing a KPI tree from the North Star 16:05 LLM agents and the future of experimentation at Zalando 18:35 New missions for a longstanding business Takeaways -Treat experimentation as a portfolio: balance confirmatory tests that protect the business with game-changing bets that can win big. -Assess risk tiers when building the roadmap so measurement rigor scales with the stakes instead of slowing every decision down. -When a launch is too new to have a success metric, pair short term A/B tests with long term holdouts from day one. -Connect every experiment to the company North Star by cascading it down to sensitive proxy metrics and controllable inputs. -Use LLM based agents as a cheap simulation layer to screen hypotheses, not as a replacement for real A/B tests. Connect with the Guest LinkedIn: https://www.linkedin.com/in/mi-tian-941b4767/ Website: https://www.zalando.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    How Zalando connects every experiment to its North Star
  2. 1d ago

    Learneo on testing the opposite of every hypothesis

    SummaryRich Liebling, senior director of engineering at Learneo, joins host Ashley Stirrup to explain the practice that came out of growing Shop It To Me from 50,000 subscribers to one million in nine months: test your hypothesis, and test its opposite. Rich covers the page where cutting text lost and adding text won, why the inverse wins more often than teams expect, and why small focused tests are the only ones where "the opposite" means anything. He also walks through translating a million subscriber goal into a target of 300 A/B tests, making a new engineer's second merge request their own experiment, and the different constraints at Course Hero, where competing team metrics were resolved with an early lifetime value model and three-day SQL analyses quietly capped testing velocity. This episode is for product managers, engineers, data scientists, and growth leaders building or scaling an experimentation program. Chapters 00:00 Cold open and introduction 01:50 Shop It To Me and the first engineering hire 02:55 300 A/B tests as the path to one million subscribers 04:00 Testing as a core value and the second merge request 07:20 Learneo, Course Hero, and two ways to get access 08:50 Modeling lifetime value to stop teams competing 11:10 Testing the opposite of the hypothesis 14:35 A portfolio of small, medium, and large tests 17:15 Multi armed bandits and seasonal traffic 20:35 The flywheel that keeps copycats behind Takeaways -Test the opposite of every hypothesis. At Shop It To Me the inverse won surprisingly often, and even when it lost it proved the variable mattered. -Keep tests small and focused. Redesign a whole page and lose, and you learn that version failed but not what to change next. -Translate a growth goal into an execution count. One million subscribers is not actionable. 300 A/B tests by year end is, and everyone can influence it. -Make experimentation part of hiring and onboarding. Every new engineer's second merge request was their own test idea. -Analysis friction sets the ceiling on testing velocity. Three days of ad hoc SQL per test quietly discourages teams from running more. Connect with the GuestLinkedIn: https://www.linkedin.com/in/richliebling/Website: https://www.learneo.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    Learneo on testing the opposite of every hypothesis
  3. 6d ago

    The four questions Early Warning asks before any A/B test

    Summary What separates a valid A/B test from an expensive guess? Priya Singhee, VP of Enterprise Analytics & Data Science at Early Warning — the bank-owned consortium that fights payment fraud and operates Zelle, which processed a trillion dollars last year — joins host Ashley Stirrup to share the experimentation playbook she built leading storefront analytics at Wayfair. She walks through the four questions to ask before launching any A/B test, why 85 to 90% of tests are supposed to fail, how pre-registration and kill criteria stop p-hacking before it starts, the pitfalls that fake wins (novelty effects, hidden heterogeneity, multiple comparisons), and how to roll out winners with gradual ramps and long-running holdouts. A practical episode for product managers, engineers, data scientists, and growth leaders building rigorous experimentation programs. Chapters 00:45 Meet Early Warning: fraud detection, Zelle, and a trillion dollars in payments 02:00 Wayfair and optimizing every step of the storefront funnel 03:05 The four questions to ask before any A/B test 05:15 Test setup best practices: hypotheses, guardrails, power, and pre-registration 07:40 Why 85 to 90% of tests fail and why that's a learning agenda 09:25 Novelty effects, hidden heterogeneity, and the multiple comparisons problem 12:55 Pre-registration, kill criteria, and stopping p-hacking 15:10 Rolling out winners: gradual ramps and long-running holdouts 17:15 Causal inference when you can't A/B test 20:05 The case for more A/B testing, not less Takeaways Run the four-question framework before any test: clean randomization, a plausible effect size for your traffic, a reversible and cheap change, and a falsifiable hypothesis. Treat A/B testing as a learning agenda: 85 to 90% of tests are supposed to fail, and a suspiciously high win rate is a red flag, not a trophy. Pre-register the full analysis plan, including hypothesis, mechanism, primary metric, exact statistical test, and subgroups, so p-hacking can't creep in when a test goes sideways. Define kill criteria and success, failure, and guardrail-dip actions before launch, with leadership sign-off, so nobody chases a loss into a fake win. Log every test and its learnings where the whole organization can see them; that reinforcement loop is what separates world-class experimentation programs. Connect with the Guest LinkedIn: https://www.linkedin.com/in/priya-singhee/ Website: https://www.earlywarning.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    The four questions Early Warning asks before any A/B test
  4. Sep 1

    How Supercell A/B tests 300 million players without breaking trust

    SummaryOn this episode of The Experimentation Edge, Ashley Stirrup talks with Shan Huang, data scientist on the central experimentation team at Supercell, the Helsinki mobile game company behind Clash of Clans, Clash Royale, Brawl Stars, Hay Day, and Boom Beach. Shan explains how a famously decentralized, creative-first company with 300 million monthly active users runs fewer than 100 A/B tests a quarter and why that restraint is deliberate, how Supercell announces experiments to players in advance and promises make-up events to keep testing fair, and how importable AI skills now let anyone at the company analyze their own experiments, making quality consistency the next big challenge. It's for product managers, data scientists, and growth leaders balancing creative conviction with experimental rigor. Chapters 00:00 Intro 01:10 About Supercell and 300 million players 02:30 The central team and a decentralized culture 04:25 Fewer than 100 tests a quarter 05:40 Sharing learnings across independent game teams 07:35 Retention as the North Star 08:35 Onboarding experiments with gems and tutorials 12:45 Telling players about A/B tests 15:15 Hypotheses and proxy metrics 20:45 AI and the future of experiment analysis Takeaways -Supercell runs fewer than 100 A/B tests a quarter for 300 million monthly players, because the goal is to become more hypothesis driven while staying creative, not to maximize volume. -In a decentralized company, a central experimentation team earns its impact by providing the platform, partnering on rigor, and making sure learnings travel across independent game teams. -Announce experiments to users before they run; Supercell's community managers tell players what is being tested and why, which turns a skeptical community into a research partner. -Promise fairness, not just transparency; players who get the worse variant always receive a make-up event later, because game players come to have fun, not to be disadvantaged. -AI-powered self-serve analysis means everyone can now run and analyze experiments, so the next challenge is making the quality of AI analysis consistent across the whole company. Connect with the GuestLinkedIn: https://www.linkedin.com/in/cnshanhuang/Website: https://supercell.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    How Supercell A/B tests 300 million players without breaking trust
  5. Aug 27

    How Clover experiments when billions of dollars flow through daily

    Summary How do you run an experimentation program when classic A/B testing is off the table? Ben Schein, Director of Product Management at Clover, joins host Ashley Stirrup to explain how a platform serving 300,000+ merchants and processing billions of dollars daily proves every feature through pilots and ground-level testing before rollout, why uncertainty and downside — not feature visibility — decide testing depth, and what his years leading product at Shake Shack taught him about turning a checkout funnel into a brand channel. This episode is for product managers, engineers, and data scientists building experimentation programs where the stakes are real. Chapters 00:00 Cold open and welcome 01:30 The Clover business model and its scale 03:45 Ben's role and the metrics that matter 07:45 Deciding what gets tested: uncertainty and downside 09:05 Why Clover can't test in production 12:45 Testing the Shake Shack checkout experience 19:35 Advice for PMs new to experimentation 23:45 Context over personalization 25:45 The future of experimentation and the human element Takeaways -At Clover's scale, testing happens before rollout: pilots and detailed go-to-market plans replace in-production A/B tests, because a merchant's work tool can never change overnight without warning. -Uncertainty and downside set the testing depth. High-risk changes like payment authorization flows earn deep, ground-level experimentation, while table stakes features like Apple Pay earn a monitored rollout. -Testing is the evidence that justifies rollout investment: if the data doesn't show a feature will succeed, the go-to-market dollars never get spent. -A checkout funnel can carry the brand. At Shake Shack, Ben's team tested prep-time expectations, fixed wrong-location orders, and used loading screens to deliver hospitality digitally. -Structure experiments for durable business value, not pass-fail verdicts: pair headline metrics with counter metrics and anchor on ground-level measurements like items per check that resist marketing noise. Connect with the Guest LinkedIn: https://www.linkedin.com/in/benschein/ Website: https://www.clover.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    How Clover experiments when billions of dollars flow through daily
  6. Aug 26

    Why JobLeads says one test won't move you, but 100 will

    SummaryWhat happens when the most logical feature you've ever built has zero impact? In this episode of The Experimentation Edge, host Ashley Stirrup, CMO of GrowthBook, sits down with Edd Saunders, product experimentation manager at JobLeads, to unpack the pizza personalization experiment that cut an ordering flow from 22 clicks to 5 and changed nothing. Edd shares the problem mapping framework he uses to move new experimenters from solution space to problem space thinking, how JobLeads grew from 0.3 to 2.8 experiments per month, and why democratizing experimentation across a whole company comes down to habit change rather than education. A practical conversation for product managers, data scientists, engineers, and growth leaders building experimentation cultures. Chapters 00:00 Introduction 01:40 Meet Edd Saunders and JobLeads 03:12 How JobLeads uses AI for prototyping 04:05 Building experimentation operations and a knowledge base 05:35 Learning over winning and compounding growth 08:10 The pizza personalization experiment 12:50 Moving from solution space to problem space 17:05 Problem mapping on a 2x2 matrix 22:20 Velocity and democratizing experimentation 24:20 AI automation for the unsexy work Takeaways -A one click reorder feature that cut a pizza ordering flow from 22 inputs to 5 had zero impact on purchases, proving that removing friction can also remove the customer's sense of control. -Exploration is part of the customer's delight; returning customers wanted to browse the menu even though they ordered the same thing every week. -Moving new experimenters from solution space to problem space thinking raises win rates and produces learnings the whole organization can use. -Problem mapping on a 2x2 matrix of evidence versus impact turns customer research into a prioritized experiment roadmap, and one validated problem can spring a whole tree of testable ideas. -Scaling experimentation from 0.3 to 2.8 tests per month is less about education and more about habit change, shared learnings, and giving non specialists the tools to launch their own experiments. Connect with the GuestLinkedIn: https://www.linkedin.com/in/eddsaunders/Website: https://www.jobleads.comRead Blog: https://www.growthbook.io/podcast/episode/1-34 SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    Why JobLeads says one test won't move you, but 100 will
  7. Aug 20

    Why The Aspen Group targets a 25% win rate

    SummaryWhat does a winning A/B test mean when your website serves 1,100 dentist owned offices? In this episode of The Experimentation Edge, host Ashley Stirrup talks with Arie Polycarpou, Senior Manager of Test and Learn at Aspen Dental, about building experimentation programs at Kohl's, Marriott, Total Wine, and now the largest company in The Aspen Group's healthcare retail portfolio. Arie explains why appointment bookings are the North Star but never the whole story, why he deliberately targets a 25% win rate and expects it to fall as the program matures, and why he applies a haircut to every stacked win before it reaches leadership. A practical conversation for product managers, engineers, data scientists, and growth leaders building test and learn programs at multi-location businesses. Chapters 00:00 introduction and Arie's path into experimentation 02:05 building programs at Kohl's, Marriott, and Total Wine 03:15 healthcare retail and 1,100 dentist owned offices 04:30 the test and learn team and 100 tests a year 07:15 building a culture of shared wins and explained losses 09:50 what a losing navigation redesign revealed 12:30 testing your way into big redesigns 13:55 the case for a 25% win rate 15:50 haircuts, holdouts, and honest math on stacked wins 17:55 metrics beyond conversion and where AI fits next Takeaways -Aspen Dental runs experimentation as healthcare retail: with 1,100 dentist owned offices, the office, not just the website visitor, is the real unit of analysis. -A mature program should target a true win rate around 25%; a 40% win rate usually signals a young program still picking off low-hanging fruit. -Apply a haircut to stacked wins: lifts depreciate as customers acclimate, and two 5% wins never add up to 10%. -Losing tests are valuable when they're designed to isolate why: test your way into big redesigns instead of shipping them whole. -Experimentation culture grows from sharing wins and explaining losses: keep the statistical rigor in the back end and the communication simple. Connect with the GuestLinkedIn: https://www.linkedin.com/in/ariepolycarpou/Website: https://www.aspendental.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    Why The Aspen Group targets a 25% win rate
  8. Aug 18

    How Principal A/B tests on customers who don't exist

    Summary What if you could A/B test on customers who don't exist before spending a single live impression? Erika Dunn, assistant director of data science at Principal Financial Group, joins host Ashley Stirrup, CMO at GrowthBook, to share how she built synthetic digital audiences entirely in house: profiles shaped by real data that rank content by likelihood of engagement, whose first live A/B test selection just beat the control. They also dig into the looping metric her team built in SQL to find where customers get stuck without heat-mapping tools, why testing gets watered down into "let's try something" at large companies, and how a center of excellence that shares wins and losses defeats the "we tried that years ago" reflex. This episode is for product managers, data scientists, marketers, and experimentation leaders, especially those working inside large, risk-averse organizations. Chapters 00:00 Cold open and welcome 01:40 Erika's path from quantitative psychology to experimentation 03:35 When testing gets watered down 06:50 Experimentation in Principal's marketing space 09:55 The looping metric that finds stuck users 12:45 Building synthetic digital audiences 15:15 The first synthetic audience A/B test wins 22:30 The PDF lesson and meeting customers on mobile 25:15 North star metrics and the right contact cadence 28:05 Where experimentation goes next with AI agents Takeaways - Synthetic digital audiences let teams rank 20 content options by predicted engagement before spending a single live impression, and Principal's first synthetic selection beat the control in a real A/B test. - A looping metric built from web behavior data can find where customers get stuck without heat-mapping tools: watch how often users cycle back to the same page within tight time windows. - Testing gets watered down when "let's try something" replaces a control group; a little pre-planning gets far more out of every experiment. - A center of excellence that shares wins and losses turns tribal knowledge into shared knowledge and stops "we tried that years ago" from killing valuable retests. - An experimentation mindset requires that people can't get punished for mistakes; give teams guardrails and a safe playground and they'll stop running the same test forever. Connect with the Guest LinkedIn: https://www.linkedin.com/in/erikadunn/ Website: https://www.principal.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    How Principal A/B tests on customers who don't exist

About

How do product teams decide what to build and what not to? The Experimentation Edge is the podcast where product, growth, and engineering leaders share how A/B testing, feature flags, and experimentation drive real business outcomes — backed by named companies and real numbers. From DoorDash's 12,000 A/B tests a year to Atlassian's experimentation-led product win to UPS's $500M experimentation team, each episode goes deep with operators running experimentation programs at scale. Hosted by Ashley Stirrup, CMO at GrowthBook and a 25-year executive in data and experimentation. For product managers, engineers, data scientists, and growth leaders at B2B tech companies who care about experimentation culture, statistical rigor, and shipping with confidence. No marketing speak. Just operators explaining what they shipped, what moved the needle, and how experimentation reshaped their teams. Topics: A/B testing, experimentation, growth experimentation, product experimentation, tech experimentation, feature flags, experimentation culture, statistical significance, marketplace experimentation, conversion rate optimization, experimentation at scale.

You Might Also Like