The Experimentation Edge

Growthbook

How do product teams decide what to build and what not to? The Experimentation Edge is the podcast where product, growth, and engineering leaders share how A/B testing, feature flags, and experimentation drive real business outcomes — backed by named companies and real numbers. From DoorDash's 12,000 A/B tests a year to Atlassian's experimentation-led product win to UPS's $500M experimentation team, each episode goes deep with operators running experimentation programs at scale. Hosted by Ashley Stirrup, CMO at GrowthBook and a 25-year executive in data and experimentation. For product managers, engineers, data scientists, and growth leaders at B2B tech companies who care about experimentation culture, statistical rigor, and shipping with confidence. No marketing speak. Just operators explaining what they shipped, what moved the needle, and how experimentation reshaped their teams. Topics: A/B testing, experimentation, growth experimentation, product experimentation, tech experimentation, feature flags, experimentation culture, statistical significance, marketplace experimentation, conversion rate optimization, experimentation at scale.

  1. 16h ago

    Why Twilio ships on signals instead of significance

    SummaryWanli Lau spent eleven years at Expedia Group, most recently running consumer product analytics on a program that shipped more than 1,500 A/B tests a year. Five months ago she moved to Twilio to lead global R&D analytics. In this episode she walks Ashley Stirrup through what carried over and what did not. The centerpiece is a test that looked like a win. At Expedia, an account sign-up takeover on the front page drove conversion sharply up, and quietly harmed return rate and engagement across every geography. That result is where Wanli's insistence on guardrail metrics and a trade-off decision framework comes from. She also explains the parts of B2B experimentation that have no B2C equivalent: choosing between user-level and account-level randomization, the interference risk when two colleagues at the same company see different pricing, and the many-to-many problem of one developer belonging to several accounts. And because Twilio's customer counts are lower than a consumer business, her teams often ship on signals and team conviction rather than waiting for statistical significance. Chapters00:00 Cold open01:02 Welcome Wanli Lau of Twilio01:23 Inside Wanli's role at Twilio04:03 Eleven years and 1,500 tests a year at Expedia04:58 Why B2B experimentation differs from B2C05:33 Cross-functional collaboration and good hypotheses06:39 Teaching teams to experiment rigorously07:48 The Expedia sign-up takeover that won on conversion11:09 Guardrail metrics and the trade-off framework14:52 Designing experiments to maximize learning17:46 Diagnosing where a feature failed19:10 North Star metrics and account-level randomization at Twilio21:40 Learning from signals instead of significance22:11 Where experimentation goes next at Twilio Takeaways- A test that wins on the primary metric can still be a loss. Expedia's sign-up takeover raised conversion and damaged return rate and engagement at the same time.- Name the primary, secondary and guardrail metrics before the test runs, not at readout. Deciding afterwards turns a result into a debate.- Conversion should be a do-no-harm guardrail for teams that do not own it, even when it is not their primary metric.- In B2B the randomization unit is a design decision. User-level bucketing risks two colleagues at one account seeing different prices, and account-level avoids that but costs sample size.- When customer counts are low, significance is often out of reach. Learning from signals and shipping on team conviction beats waiting for a number that will never arrive. Connect with the GuestWanli Lau LinkedIn: https://www.linkedin.com/in/wanlilauAbout the guest: https://www.growthbook.io/podcast/guests/wanli-lauEpisode page: https://www.growthbook.io/podcast/episode/1-44Company Website: https://www.twilio.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to growthbook.io

    Why Twilio ships on signals instead of significance
  2. 2d ago

    Why GoPro stopped judging A/B tests by win rate

    Will Guyeskey is Director of Digital Product at GoPro, where his team owns the e-commerce side of gopro.com. Before GoPro he ran personalization at Gap and cut his teeth at Brooks Bell, testing for brands like Barnes & Noble, Under Armour and Ralph Lauren. In this episode of The Experimentation Edge, he tells Ashley Stirrup why knowing your customer is the through line of every good test program. Will shares the Barnes & Noble order confirmation test he was sure would lose, and why the same idea never worked for any other client. He walks through a recent GoPro Mission launch test that asked whether a step-by-step configurator adds too much friction, and what a flat result revealed about high consideration buyers. He also explains why GoPro shares interim readouts across the company, why win rate makes a poor North Star for an experimentation program, and how his team plans to use AI for speed without outrunning its own learnings. Chapters00:00 Intro00:52 Will's role running e-commerce at GoPro01:43 Learning A/B testing across retail at Brooks Bell05:24 How GoPro runs one to three tests a month06:35 Sharing learnings and interim readouts across teams09:20 The Barnes & Noble order confirmation win13:02 Why the win did not transfer to other clients14:41 Testing friction on the GoPro Mission configurator18:51 Why win rate is the wrong North Star22:19 How AI will shape experimentation at GoPro Takeaways- A winning idea rarely travels. The Barnes & Noble recommendation module worked because of that audience's low order values and reading habits, and it failed for every other client that tried it.- Design every test so it teaches you something whether it wins, loses or ends flat. Losing tests are jet fuel when the learning is built in.- A flat result is still an answer. GoPro's configurator test showed that buyers of high consideration products accept extra steps when each choice adds value.- Share interim readouts across the company, and use them to show how volatile results are before a test reaches statistical significance.- AI can speed up building and running experiments, but a team that runs more tests than it can learn from is not getting better. Connect with the GuestWill Guyeskey LinkedIn: https://www.linkedin.com/in/willguyeskey/About the guest: https://www.growthbook.io/podcast/guests/will-guyeskeyEpisode page: https://www.growthbook.io/podcast/episode/1-43Company Website: https://gopro.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to growthbook.io

    Why GoPro stopped judging A/B tests by win rate
  3. Sep 17

    What Samsung learned bringing B2C rigor to B2B

    SummaryWhat changes when you take a mature B2C experimentation practice and apply it to a B2B storefront? Anuradha Tempe, Lead Product Manager at Samsung Electronics America, joins host Ashley Stirrup to share what she learned managing both sides of Samsung.com. She explains why B2B is not a sidekick to B2C, how a single bulk order can push a test into a false positive unless you normalize the data, and why not everything deserves an A/B test. She walks through the add-on experiment that nearly doubled attach sales once her team realized business buyers decide on mobile and purchase on desktop, and the Buy Now test that lost because B2B buyers value clarity over faster conversions. The conversation closes with her approach to North Star and guardrail metrics and the AI copilot she built to draft A/B test plans with a human still in the loop. A practical episode for product managers, engineers, data scientists, and growth leaders running experimentation across different customer types. Chapters00:45 Meet Anuradha Tempe: from chip design to leading Samsung e-commerce02:15 Why B2B is not a sidekick to B2C03:30 Not everything needs an A/B test05:15 Spreading experimentation practice across a global conglomerate07:00 Bringing B2C rigor to a fast-paced B2B team08:45 The add-on experiment that nearly doubled attach sales11:15 Buyers decide on mobile and purchase on desktop16:30 The Buy Now test that lost21:30 North Star metrics, guardrails, and the EPP discount fix26:05 An AI copilot for A/B test planning Takeaways- Treat B2B as its own customer base with its own testing discipline; a single bulk order on one day can inflate a B2B test into a false positive unless order data is normalized before results are read.- Not everything needs an A/B test; route lower-risk changes through UAT feedback or pre/post comparisons and reserve full experiments for features where being wrong is expensive.- Map where the decision happens, not just where the purchase happens; Samsung's business buyers decide on mobile and buy on desktop, and surfacing add-ons on mobile nearly doubled attach sales.- B2B buyers value clarity over faster conversions; a Buy Now button earlier in the flow confused bulk purchasers because returns and cancellations on large orders are costly.- Align on the North Star before building, whether it is revenue, engagement, NPS, or fewer support tickets, and set guardrails so an engagement feature can never quietly drag sales down. Connect with the GuestAnuradha Tempe LinkedIn: https://www.linkedin.com/in/anuradha-tempe/About the guest: https://www.growthbook.io/podcast/guests/anuradha-tempeEpisode page: https://www.growthbook.io/podcast/episode/1-42Company Website: https://www.samsung.com/us/business/ SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    What Samsung learned bringing B2C rigor to B2B
  4. Sep 15

    Even a loss is a win: Charlie Health's approach to experiments

    SummaryOn this episode of The Experimentation Edge, Ashley Stirrup talks with Joe Yevoli, Director of Growth at Charlie Health, a virtual intensive outpatient program that sits between weekly therapy and hospitalization. Joe explains how removing a page from Charlie Health's intake form produced a winning test that created a new bottleneck further down the funnel, why Teachers Pay Teachers cut about 80% of its market for one feature and roughly quadrupled retention, and how premortems that ask "what went wrong?" before launch make dissent safe and surface safeguards that avoid catastrophe. He also shares the two questions he asks before any experiment, why more top-of-funnel traffic always dents conversion rate, and how AI is flattening the product pod in ways that speed teams up and put them at risk. It's for growth leaders, product managers, and experimentation teams who want to learn as much from a loss as from a win. Chapters00:00 Intro01:00 About Charlie Health02:05 Joe's role across performance, lifecycle, and experimentation03:30 Building experimentation rigor across teams04:40 Even a loss is a win05:15 The form page test that created a new bottleneck08:50 Teachers Pay Teachers and the Easel lesson16:30 Designing experiments that teach you something when they lose21:00 Premortems for high-risk experiments24:30 AI is flattening the org Takeaways- A winning test is a data point, not a finish line; Charlie Health's form page removal won on top-of-funnel metrics and still exposed a downstream bottleneck that became the next experiment.- More traffic entering the funnel means conversion rate goes down, even when more people reach the bottom; treat it as a law of physics and plan for it before you read the results.- Build it for everyone and you build it for nobody; Teachers Pay Teachers cut roughly 80% of the market for Easel, repositioned it around one job, and roughly quadrupled retention.- Run a premortem before any large, risky experiment: ask the room "it was a massive failure, what went wrong?", have everyone write and share, and build the safeguards while there is still time.- Before launching, check the quantitative and qualitative data behind the hypothesis, map the full funnel, and measure toward the real business outcome, so a loss still leaves you with a next step. Connect with the GuestJoe Yevoli LinkedIn: https://www.linkedin.com/in/joeyevoli/About the guest: https://www.growthbook.io/podcast/guests/joe-yevoliEpisode page: https://www.growthbook.io/podcast/episode/1-41Company Website: https://www.charliehealth.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    Even a loss is a win: Charlie Health's approach to experiments
  5. Sep 10

    Realtor.com on using your AI as a junior data scientist

    SummaryWhat do you do when your biggest experiment win turns out to be a loss? Whitney Perez, Director of Product Management at Realtor.com, joins host Ashley Stirrup to share the checkout bundling test that posted a 300% attach rate and still lost revenue, the 30/30/30 rule she uses to set expectations for a new experimentation team, and how AI is turning an English major into an aspirational data scientist. This episode is for product managers, engineers, and data scientists building experimentation programs from the ground up. Chapters00:00 Cold open and welcome01:10 From growth hacker to Realtor.com02:50 Three foundations for a new experimentation team04:20 The 30/30/30 rule05:05 The 300% bundling win that lost revenue07:10 You don't need a stats degree to experiment09:20 Cascading North Star metrics12:50 Do the homework before the experiment13:50 The wishlist: instrumentation, embedded knowledge, culture15:50 AI as an aspirational data scientist18:45 Keeping a human in the loop Takeaways- A winning decision metric is not enough. Realtor.com's bundling test hit a 300% attach rate, but funnel fallout from the extra step made it a net revenue loser. Set secondary metrics and their thresholds before launch.- Expect the 30/30/30 rule: roughly a third of tests win, a third are inconclusive, and a third lose. The math is the math, and the losers carry most of the learning.- Start a new team on foundations: what a clean test and an A/A test look like, which surfaces should not be tested, and a peer review program that lets people graduate to more complex experiments.- You don't need a stats background to run good experiments. Teach the simplest definition of a good test, then let people learn by doing.- AI can make anyone an aspirational data scientist for analyzing results and spotting opportunities, but it can be confidently wrong. Keep a human in the loop and sanity check output the way you'd peek at a freshly launched test. Connect with the GuestWhitney Perez LinkedIn: https://www.linkedin.com/in/whitneykperez/About the guest: https://www.growthbook.io/podcast/guests/whitney-perezEpisode page: https://www.growthbook.io/podcast/episode/1-40Company Website: https://www.realtor.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    Realtor.com on using your AI as a junior data scientist
  6. Sep 9

    How Zalando connects every experiment to its North Star

    SummaryHow do you keep 1,000 experiments a year pointed at one North Star? In this episode of The Experimentation Edge, host Ashley Stirrup talks with Mi Tian, Head of Applied Science at Zalando, about running experimentation inside a central economics org that reports to the CFO. Mi shares how Zalando balances safe confirmatory tests with game-changing bets, how the team measured the discovery feeds homepage launch when success had no established metric, and how a KPI tree cascades the company North Star down to the controllable inputs teams ship every day. She also looks ahead to LLM-based agents as a simulation layer for screening hypotheses. A practical conversation for anyone building an experimentation program that wants both rigor and ambition. Chapters00:45 About Zalando and its marketplace model02:00 Mi's path from engineering to experimentation03:10 Economists and data scientists in one decision-making org04:15 Running over 1,000 experiments a year05:55 What makes an experiment high risk07:10 Sharing learnings through standardization and champions09:10 The discovery feeds launch and its measurement plan11:15 Balancing the experimentation portfolio12:35 Growing a KPI tree from the North Star16:05 LLM agents and the future of experimentation at Zalando18:35 New missions for a longstanding business Takeaways- Treat experimentation as a portfolio: balance confirmatory tests that protect the business with game-changing bets that can win big.- Assess risk tiers when building the roadmap so measurement rigor scales with the stakes instead of slowing every decision down.- When a launch is too new to have a success metric, pair short term A/B tests with long term holdouts from day one.- Connect every experiment to the company North Star by cascading it down to sensitive proxy metrics and controllable inputs.- Use LLM based agents as a cheap simulation layer to screen hypotheses, not as a replacement for real A/B tests. Connect with the GuestMi Tian LinkedIn: https://www.linkedin.com/in/mi-tian-941b4767/About the guest: https://www.growthbook.io/podcast/guests/mi-tianEpisode page: https://www.growthbook.io/podcast/episode/1-39Company Website: https://www.zalando.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    How Zalando connects every experiment to its North Star
  7. Sep 8

    Learneo on testing the opposite of every hypothesis

    SummaryRich Liebling, senior director of engineering at Learneo, joins host Ashley Stirrup to explain the practice that came out of growing Shop It To Me from 50,000 subscribers to one million in nine months: test your hypothesis, and test its opposite. Rich covers the page where cutting text lost and adding text won, why the inverse wins more often than teams expect, and why small focused tests are the only ones where "the opposite" means anything. He also walks through translating a million subscriber goal into a target of 300 A/B tests, making a new engineer's second merge request their own experiment, and the different constraints at Course Hero, where competing team metrics were resolved with an early lifetime value model and three-day SQL analyses quietly capped testing velocity. This episode is for product managers, engineers, data scientists, and growth leaders building or scaling an experimentation program. Chapters00:00 Cold open and introduction01:50 Shop It To Me and the first engineering hire02:55 300 A/B tests as the path to one million subscribers04:00 Testing as a core value and the second merge request07:20 Learneo, Course Hero, and two ways to get access08:50 Modeling lifetime value to stop teams competing11:10 Testing the opposite of the hypothesis14:35 A portfolio of small, medium, and large tests17:15 Multi armed bandits and seasonal traffic20:35 The flywheel that keeps copycats behind Takeaways- Test the opposite of every hypothesis. At Shop It To Me the inverse won surprisingly often, and even when it lost it proved the variable mattered.- Keep tests small and focused. Redesign a whole page and lose, and you learn that version failed but not what to change next.- Translate a growth goal into an execution count. One million subscribers is not actionable. 300 A/B tests by year end is, and everyone can influence it.- Make experimentation part of hiring and onboarding. Every new engineer's second merge request was their own test idea.- Analysis friction sets the ceiling on testing velocity. Three days of ad hoc SQL per test quietly discourages teams from running more. Connect with the GuestRich Liebling LinkedIn: https://www.linkedin.com/in/richliebling/About the guest: https://www.growthbook.io/podcast/guests/rich-lieblingEpisode page: https://www.growthbook.io/podcast/episode/1-38Company Website: https://www.learneo.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    Learneo on testing the opposite of every hypothesis
  8. Sep 3

    The four questions Early Warning asks before any A/B test

    SummaryWhat separates a valid A/B test from an expensive guess? Priya Singhee, VP of Enterprise Analytics & Data Science at Early Warning — the bank-owned consortium that fights payment fraud and operates Zelle, which processed a trillion dollars last year — joins host Ashley Stirrup to share the experimentation playbook she built leading storefront analytics at Wayfair. She walks through the four questions to ask before launching any A/B test, why 85 to 90% of tests are supposed to fail, how pre-registration and kill criteria stop p-hacking before it starts, the pitfalls that fake wins (novelty effects, hidden heterogeneity, multiple comparisons), and how to roll out winners with gradual ramps and long-running holdouts. A practical episode for product managers, engineers, data scientists, and growth leaders building rigorous experimentation programs. Chapters00:45 Meet Early Warning: fraud detection, Zelle, and a trillion dollars in payments02:00 Wayfair and optimizing every step of the storefront funnel03:05 The four questions to ask before any A/B test05:15 Test setup best practices: hypotheses, guardrails, power, and pre-registration07:40 Why 85 to 90% of tests fail and why that's a learning agenda09:25 Novelty effects, hidden heterogeneity, and the multiple comparisons problem12:55 Pre-registration, kill criteria, and stopping p-hacking15:10 Rolling out winners: gradual ramps and long-running holdouts17:15 Causal inference when you can't A/B test20:05 The case for more A/B testing, not less Takeaways- Run the four-question framework before any test: clean randomization, a plausible effect size for your traffic, a reversible and cheap change, and a falsifiable hypothesis.- Treat A/B testing as a learning agenda: 85 to 90% of tests are supposed to fail, and a suspiciously high win rate is a red flag, not a trophy.- Pre-register the full analysis plan, including hypothesis, mechanism, primary metric, exact statistical test, and subgroups, so p-hacking can't creep in when a test goes sideways.- Define kill criteria and success, failure, and guardrail-dip actions before launch, with leadership sign-off, so nobody chases a loss into a fake win.- Log every test and its learnings where the whole organization can see them; that reinforcement loop is what separates world-class experimentation programs. Connect with the GuestPriya Singhee LinkedIn: https://www.linkedin.com/in/priya-singhee/About the guest: https://www.growthbook.io/podcast/guests/priya-singheeEpisode page: https://www.growthbook.io/podcast/episode/1-37Company Website: https://www.earlywarning.com SponsorGrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide. Go to http://growthbook.io

    The four questions Early Warning asks before any A/B test
5
out of 5
10 Ratings

About

How do product teams decide what to build and what not to? The Experimentation Edge is the podcast where product, growth, and engineering leaders share how A/B testing, feature flags, and experimentation drive real business outcomes — backed by named companies and real numbers. From DoorDash's 12,000 A/B tests a year to Atlassian's experimentation-led product win to UPS's $500M experimentation team, each episode goes deep with operators running experimentation programs at scale. Hosted by Ashley Stirrup, CMO at GrowthBook and a 25-year executive in data and experimentation. For product managers, engineers, data scientists, and growth leaders at B2B tech companies who care about experimentation culture, statistical rigor, and shipping with confidence. No marketing speak. Just operators explaining what they shipped, what moved the needle, and how experimentation reshaped their teams. Topics: A/B testing, experimentation, growth experimentation, product experimentation, tech experimentation, feature flags, experimentation culture, statistical significance, marketplace experimentation, conversion rate optimization, experimentation at scale.

You Might Also Like