The Experimentation Edge

Growthbook

How do product teams decide what to build and what not to? The Experimentation Edge is the podcast where product, growth, and engineering leaders share how A/B testing, feature flags, and experimentation drive real business outcomes — backed by named companies and real numbers. From DoorDash's 12,000 A/B tests a year to Atlassian's experimentation-led product win to UPS's $500M experimentation team, each episode goes deep with operators running experimentation programs at scale. Hosted by Ashley Stirrup, CMO at GrowthBook and a 25-year executive in data and experimentation. For product managers, engineers, data scientists, and growth leaders at B2B tech companies who care about experimentation culture, statistical rigor, and shipping with confidence. No marketing speak. Just operators explaining what they shipped, what moved the needle, and how experimentation reshaped their teams. Topics: A/B testing, experimentation, growth experimentation, product experimentation, tech experimentation, feature flags, experimentation culture, statistical significance, marketplace experimentation, conversion rate optimization, experimentation at scale.

  1. 15h ago

    How MilliporeSigma lifted add to cart with a smarter search

    SummaryDorothy Crepin runs the A/B testing program at MilliporeSigma, the life sciences business of Merck KGaA, where one global catalog has to serve three very different visitors: scientists hunting for a specific instrument, procurement buyers working off a list someone else wrote, and students doing research with no purchasing power at all. Add a regulated industry where what you can sell changes country by country, and personas stop being a marketing exercise. She walks Ashley Stirrup through the test that changed how her team thinks about discovery. MilliporeSigma rebuilt its on-site type ahead so it suggests a search term instead of pushing a matching product, and saw double digit increases in feature usage and add to cart. From there Dorothy makes the case for measuring the action a page is actually responsible for rather than revenue alone, for pairing A/B tests with a user testing ambassador panel to get at the why, and for sending product managers back to one question before anything gets built: what problem are you solving for the customer? Chapters00:00 Revenue is an outcome, not the metric01:26 Welcome and what MilliporeSigma does02:22 The three audiences on a B2B science site03:53 What an A/B testing analyst coordinates04:20 Running experiments across product pods05:31 Building leadership buy in for testing08:08 Where A/B testing ends and user testing begins08:52 The type ahead search test that lifted add to cart11:30 How to frame a test with a product manager13:34 Why testing the opposite pays off15:26 Personas, personalization and what AI changes18:57 Building trust across competing teams Takeaways- MilliporeSigma changed its on site search to suggest search terms instead of products and saw double digit gains in feature usage and add to cart.- Revenue is an outcome, so Dorothy's team measures the next action each page is responsible for, like add to cart on a product detail page.- A/B testing tells you what happened. MilliporeSigma's user testing ambassador panel supplies the why behind the click.- Every test starts with one question to the product manager: what problem are you solving for the customer, not what is your hypothesis.- Dorothy is building personas into agents so browsers, not just buyers, can be tested for a directional read before an experiment ships. Connect with the GuestDorothy Crepin LinkedIn: https://www.linkedin.com/in/dorothycrepin/About the guest: https://www.growthbook.io/podcast/guests/dorothy-crepinEpisode page: https://www.growthbook.io/podcast/episode/1-47Company Website: https://www.sigmaaldrich.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to growthbook.io

    How MilliporeSigma lifted add to cart with a smarter search
  2. 6d ago

    PayPal's $180 million experimentation win

    SummaryGaurav Sethi joins Ashley Stirrup on The Experimentation Edge to explain how PayPal's experimentation program went from 700 to 800 tests a year with four week readouts to 2,587 experiments a year and roughly $180 million in measured impact. Gaurav inherited a reported win rate of 55% to 60%, four times the industry average, and traced it to experiments logging assignment data instead of exposure data, carrying 25% to 30% dilution. The conversation covers the exposure event his team introduced, the instrumentation and metric standards that cut readouts to 24 hours, a carousel test that a multi armed bandit resolved in 51 days instead of a projected 700, and why cost avoidance from losing experiments belongs in the ROI number. It closes on what changes when the thing you are testing is an AI agent rather than a button. Useful for product managers, engineers, data scientists and growth leaders building or defending an experimentation program at scale. Chapters00:00 Cold open00:54 Welcome Gaurav Sethi of PayPal01:30 Elmo, PayPal's homegrown experimentation platform03:53 $180 million in revenue impact and cost avoidance04:28 The win rate that was too good to be true06:29 Instrumentation standards and the exposure event08:56 The carousel test: 700 days down to 5111:29 Designing experiments so every result teaches you13:45 Building the platform is only half the job15:07 Education, office hours and executive support18:28 Exposure events and joining transactional data24:14 AI for experimentation, and experimentation for AI Takeaways- A win rate of 55% to 60% against an industry average of 11% to 14% was a tracking problem, not a performance one: assignment data carried 25% to 30% dilution and reached significance on users who never saw the test.- The exposure event fixed it. Fire an event as close to render as possible so the platform knows exactly when a user entered the experiment, instead of logging on page load.- Standardized instrumentation and canonical metric definitions are what let the analysis pipeline run without human cleanup, taking experiment readouts from four weeks to 24 hours.- A six variant carousel test projected at 700 days was resolved by a multi armed bandit in 51 days, with a winner worth almost five basis points.- About half of PayPal's $180 million impact in 2025 was cost avoidance from features that tested badly and never shipped, which is why Gaurav argues there are no losing experiments. Connect with the GuestGaurav Sethi LinkedIn: https://www.linkedin.com/in/gauravsethi22About the guest: https://www.growthbook.io/podcast/guests/gaurav-sethiEpisode page: https://www.growthbook.io/podcast/episode/1-46Company Website: https://www.paypal.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to growthbook.io

    PayPal's $180 million experimentation win
  3. Sep 29

    ServiceNow's Customer Zero approach to AI experimentation

    SummaryServiceNow runs its own platform on itself. Ashraf Karim, Senior Vice President, Connected Customer Experience, owns everything that happens after a customer buys, and her team is Customer Zero, testing every product and capability internally before it ever reaches a customer. She joins Ashley Stirrup to unpack what that vantage point teaches you about experimenting with AI. The headline lesson is not a technical one. ServiceNow built agentic skills that worked, and users abandoned them anyway, because a task that ran five to fifteen minutes broke every expectation people had about what a machine should feel like. The fix was not a faster model. It was telling the user what the AI was doing while it did it. Ashraf also gets into the harder measurement problem underneath all of it: how do you evaluate a nondeterministic system? ServiceNow's answer is an internal LLM judge scored against a golden data set, plus a customer effort score that catches what CSAT reports too late. Chapters00:00 What customers actually care about01:14 Meet Ashraf Karim of ServiceNow01:39 Owning the entire post-sale experience02:08 Why ServiceNow runs as its own Customer Zero03:02 Lessons from Google, PayPal and Verizon07:11 The agentic skills users kept abandoning09:29 Latency, expectations and the feedback loop fix12:29 Users wanted the outcome, not the process14:49 Designing experiments that can afford to lose18:13 Using an LLM judge on nondeterministic models21:48 Customer effort score as a guardrail metric25:03 Where experimentation at ServiceNow goes next Takeaways- ServiceNow is its own Customer Zero, running the platform on itself and testing every feature internally before customers see it.- The agentic skills worked. Users still abandoned them, because a five to fifteen minute wait broke their expectation of what AI should feel like.- The fix was a feedback loop that narrates what the AI is doing in the background, which buys the patience a long task needs.- Users did not want the agentified version of the human process. They wanted the outcome and the next best action, not the steps.- Nondeterministic output needs its own evaluation layer, so ServiceNow built an LLM judge that scores against a golden data set before anything reaches a customer. Connect with the GuestAshraf Karim LinkedIn: Ashraf Karim on LinkedInAbout the guest: https://www.growthbook.io/podcast/guests/ashraf-karimEpisode page: https://www.growthbook.io/podcast/episode/1-45Company Website: ServiceNow SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to growthbook.io

    ServiceNow's Customer Zero approach to AI experimentation
  4. Sep 24

    Why Twilio ships on signals instead of significance

    SummaryWanli Lau spent eleven years at Expedia Group, most recently running consumer product analytics on a program that shipped more than 1,500 A/B tests a year. Five months ago she moved to Twilio to lead global R&D analytics. In this episode she walks Ashley Stirrup through what carried over and what did not. The centerpiece is a test that looked like a win. At Expedia, an account sign-up takeover on the front page drove conversion sharply up, and quietly harmed return rate and engagement across every geography. That result is where Wanli's insistence on guardrail metrics and a trade-off decision framework comes from. She also explains the parts of B2B experimentation that have no B2C equivalent: choosing between user-level and account-level randomization, the interference risk when two colleagues at the same company see different pricing, and the many-to-many problem of one developer belonging to several accounts. And because Twilio's customer counts are lower than a consumer business, her teams often ship on signals and team conviction rather than waiting for statistical significance. Chapters00:00 Cold open01:02 Welcome Wanli Lau of Twilio01:23 Inside Wanli's role at Twilio04:03 Eleven years and 1,500 tests a year at Expedia04:58 Why B2B experimentation differs from B2C05:33 Cross-functional collaboration and good hypotheses06:39 Teaching teams to experiment rigorously07:48 The Expedia sign-up takeover that won on conversion11:09 Guardrail metrics and the trade-off framework14:52 Designing experiments to maximize learning17:46 Diagnosing where a feature failed19:10 North Star metrics and account-level randomization at Twilio21:40 Learning from signals instead of significance22:11 Where experimentation goes next at Twilio Takeaways- A test that wins on the primary metric can still be a loss. Expedia's sign-up takeover raised conversion and damaged return rate and engagement at the same time.- Name the primary, secondary and guardrail metrics before the test runs, not at readout. Deciding afterwards turns a result into a debate.- Conversion should be a do-no-harm guardrail for teams that do not own it, even when it is not their primary metric.- In B2B the randomization unit is a design decision. User-level bucketing risks two colleagues at one account seeing different prices, and account-level avoids that but costs sample size.- When customer counts are low, significance is often out of reach. Learning from signals and shipping on team conviction beats waiting for a number that will never arrive. Connect with the GuestWanli Lau LinkedIn: https://www.linkedin.com/in/wanlilauAbout the guest: https://www.growthbook.io/podcast/guests/wanli-lauEpisode page: https://www.growthbook.io/podcast/episode/1-44Company Website: https://www.twilio.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to growthbook.io

    Why Twilio ships on signals instead of significance
  5. Sep 22

    Why GoPro stopped judging A/B tests by win rate

    Will Guyeskey is Director of Digital Product at GoPro, where his team owns the e-commerce side of gopro.com. Before GoPro he ran personalization at Gap and cut his teeth at Brooks Bell, testing for brands like Barnes & Noble, Under Armour and Ralph Lauren. In this episode of The Experimentation Edge, he tells Ashley Stirrup why knowing your customer is the through line of every good test program. Will shares the Barnes & Noble order confirmation test he was sure would lose, and why the same idea never worked for any other client. He walks through a recent GoPro Mission launch test that asked whether a step-by-step configurator adds too much friction, and what a flat result revealed about high consideration buyers. He also explains why GoPro shares interim readouts across the company, why win rate makes a poor North Star for an experimentation program, and how his team plans to use AI for speed without outrunning its own learnings. Chapters00:00 Intro00:52 Will's role running e-commerce at GoPro01:43 Learning A/B testing across retail at Brooks Bell05:24 How GoPro runs one to three tests a month06:35 Sharing learnings and interim readouts across teams09:20 The Barnes & Noble order confirmation win13:02 Why the win did not transfer to other clients14:41 Testing friction on the GoPro Mission configurator18:51 Why win rate is the wrong North Star22:19 How AI will shape experimentation at GoPro Takeaways- A winning idea rarely travels. The Barnes & Noble recommendation module worked because of that audience's low order values and reading habits, and it failed for every other client that tried it.- Design every test so it teaches you something whether it wins, loses or ends flat. Losing tests are jet fuel when the learning is built in.- A flat result is still an answer. GoPro's configurator test showed that buyers of high consideration products accept extra steps when each choice adds value.- Share interim readouts across the company, and use them to show how volatile results are before a test reaches statistical significance.- AI can speed up building and running experiments, but a team that runs more tests than it can learn from is not getting better. Connect with the GuestWill Guyeskey LinkedIn: https://www.linkedin.com/in/willguyeskey/About the guest: https://www.growthbook.io/podcast/guests/will-guyeskeyEpisode page: https://www.growthbook.io/podcast/episode/1-43Company Website: https://gopro.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to growthbook.io

    Why GoPro stopped judging A/B tests by win rate
  6. Sep 17

    What Samsung learned bringing B2C rigor to B2B

    SummaryWhat changes when you take a mature B2C experimentation practice and apply it to a B2B storefront? Anuradha Tempe, Lead Product Manager at Samsung Electronics America, joins host Ashley Stirrup to share what she learned managing both sides of Samsung.com. She explains why B2B is not a sidekick to B2C, how a single bulk order can push a test into a false positive unless you normalize the data, and why not everything deserves an A/B test. She walks through the add-on experiment that nearly doubled attach sales once her team realized business buyers decide on mobile and purchase on desktop, and the Buy Now test that lost because B2B buyers value clarity over faster conversions. The conversation closes with her approach to North Star and guardrail metrics and the AI copilot she built to draft A/B test plans with a human still in the loop. A practical episode for product managers, engineers, data scientists, and growth leaders running experimentation across different customer types. Chapters00:45 Meet Anuradha Tempe: from chip design to leading Samsung e-commerce02:15 Why B2B is not a sidekick to B2C03:30 Not everything needs an A/B test05:15 Spreading experimentation practice across a global conglomerate07:00 Bringing B2C rigor to a fast-paced B2B team08:45 The add-on experiment that nearly doubled attach sales11:15 Buyers decide on mobile and purchase on desktop16:30 The Buy Now test that lost21:30 North Star metrics, guardrails, and the EPP discount fix26:05 An AI copilot for A/B test planning Takeaways- Treat B2B as its own customer base with its own testing discipline; a single bulk order on one day can inflate a B2B test into a false positive unless order data is normalized before results are read.- Not everything needs an A/B test; route lower-risk changes through UAT feedback or pre/post comparisons and reserve full experiments for features where being wrong is expensive.- Map where the decision happens, not just where the purchase happens; Samsung's business buyers decide on mobile and buy on desktop, and surfacing add-ons on mobile nearly doubled attach sales.- B2B buyers value clarity over faster conversions; a Buy Now button earlier in the flow confused bulk purchasers because returns and cancellations on large orders are costly.- Align on the North Star before building, whether it is revenue, engagement, NPS, or fewer support tickets, and set guardrails so an engagement feature can never quietly drag sales down. Connect with the GuestAnuradha Tempe LinkedIn: https://www.linkedin.com/in/anuradha-tempe/About the guest: https://www.growthbook.io/podcast/guests/anuradha-tempeEpisode page: https://www.growthbook.io/podcast/episode/1-42Company Website: https://www.samsung.com/us/business/ SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    What Samsung learned bringing B2C rigor to B2B
  7. Sep 15

    Even a loss is a win: Charlie Health's approach to experiments

    SummaryOn this episode of The Experimentation Edge, Ashley Stirrup talks with Joe Yevoli, Director of Growth at Charlie Health, a virtual intensive outpatient program that sits between weekly therapy and hospitalization. Joe explains how removing a page from Charlie Health's intake form produced a winning test that created a new bottleneck further down the funnel, why Teachers Pay Teachers cut about 80% of its market for one feature and roughly quadrupled retention, and how premortems that ask "what went wrong?" before launch make dissent safe and surface safeguards that avoid catastrophe. He also shares the two questions he asks before any experiment, why more top-of-funnel traffic always dents conversion rate, and how AI is flattening the product pod in ways that speed teams up and put them at risk. It's for growth leaders, product managers, and experimentation teams who want to learn as much from a loss as from a win. Chapters00:00 Intro01:00 About Charlie Health02:05 Joe's role across performance, lifecycle, and experimentation03:30 Building experimentation rigor across teams04:40 Even a loss is a win05:15 The form page test that created a new bottleneck08:50 Teachers Pay Teachers and the Easel lesson16:30 Designing experiments that teach you something when they lose21:00 Premortems for high-risk experiments24:30 AI is flattening the org Takeaways- A winning test is a data point, not a finish line; Charlie Health's form page removal won on top-of-funnel metrics and still exposed a downstream bottleneck that became the next experiment.- More traffic entering the funnel means conversion rate goes down, even when more people reach the bottom; treat it as a law of physics and plan for it before you read the results.- Build it for everyone and you build it for nobody; Teachers Pay Teachers cut roughly 80% of the market for Easel, repositioned it around one job, and roughly quadrupled retention.- Run a premortem before any large, risky experiment: ask the room "it was a massive failure, what went wrong?", have everyone write and share, and build the safeguards while there is still time.- Before launching, check the quantitative and qualitative data behind the hypothesis, map the full funnel, and measure toward the real business outcome, so a loss still leaves you with a next step. Connect with the GuestJoe Yevoli LinkedIn: https://www.linkedin.com/in/joeyevoli/About the guest: https://www.growthbook.io/podcast/guests/joe-yevoliEpisode page: https://www.growthbook.io/podcast/episode/1-41Company Website: https://www.charliehealth.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    Even a loss is a win: Charlie Health's approach to experiments
  8. Sep 10

    Realtor.com on using your AI as a junior data scientist

    SummaryWhat do you do when your biggest experiment win turns out to be a loss? Whitney Perez, Director of Product Management at Realtor.com, joins host Ashley Stirrup to share the checkout bundling test that posted a 300% attach rate and still lost revenue, the 30/30/30 rule she uses to set expectations for a new experimentation team, and how AI is turning an English major into an aspirational data scientist. This episode is for product managers, engineers, and data scientists building experimentation programs from the ground up. Chapters00:00 Cold open and welcome01:10 From growth hacker to Realtor.com02:50 Three foundations for a new experimentation team04:20 The 30/30/30 rule05:05 The 300% bundling win that lost revenue07:10 You don't need a stats degree to experiment09:20 Cascading North Star metrics12:50 Do the homework before the experiment13:50 The wishlist: instrumentation, embedded knowledge, culture15:50 AI as an aspirational data scientist18:45 Keeping a human in the loop Takeaways- A winning decision metric is not enough. Realtor.com's bundling test hit a 300% attach rate, but funnel fallout from the extra step made it a net revenue loser. Set secondary metrics and their thresholds before launch.- Expect the 30/30/30 rule: roughly a third of tests win, a third are inconclusive, and a third lose. The math is the math, and the losers carry most of the learning.- Start a new team on foundations: what a clean test and an A/A test look like, which surfaces should not be tested, and a peer review program that lets people graduate to more complex experiments.- You don't need a stats background to run good experiments. Teach the simplest definition of a good test, then let people learn by doing.- AI can make anyone an aspirational data scientist for analyzing results and spotting opportunities, but it can be confidently wrong. Keep a human in the loop and sanity check output the way you'd peek at a freshly launched test. Connect with the GuestWhitney Perez LinkedIn: https://www.linkedin.com/in/whitneykperez/About the guest: https://www.growthbook.io/podcast/guests/whitney-perezEpisode page: https://www.growthbook.io/podcast/episode/1-40Company Website: https://www.realtor.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    Realtor.com on using your AI as a junior data scientist
5
out of 5
10 Ratings

About

How do product teams decide what to build and what not to? The Experimentation Edge is the podcast where product, growth, and engineering leaders share how A/B testing, feature flags, and experimentation drive real business outcomes — backed by named companies and real numbers. From DoorDash's 12,000 A/B tests a year to Atlassian's experimentation-led product win to UPS's $500M experimentation team, each episode goes deep with operators running experimentation programs at scale. Hosted by Ashley Stirrup, CMO at GrowthBook and a 25-year executive in data and experimentation. For product managers, engineers, data scientists, and growth leaders at B2B tech companies who care about experimentation culture, statistical rigor, and shipping with confidence. No marketing speak. Just operators explaining what they shipped, what moved the needle, and how experimentation reshaped their teams. Topics: A/B testing, experimentation, growth experimentation, product experimentation, tech experimentation, feature flags, experimentation culture, statistical significance, marketplace experimentation, conversion rate optimization, experimentation at scale.

You Might Also Like