Data Breakthroughs: Solving Real-World Data Challenges

Lior Barak - Cooking Data

A podcast where data experts solve real-world operational challenges submitted by listeners. Each episode tackles a fresh problem, delivering actionable solutions, key insights, and implementation steps to help data professionals overcome barriers and create business value. impactoperations.substack.com

  1. Sep 16

    Forty Charts, No Decision

    Data Breakthroughs · Season 2, Episode 2 Runtime: [[57]] minutes · Category: Data Analysis & Reporting A note before you start: this is one of the last episodes in this format — and I’m asking for your view on it below. Why this one is worth an hour There’s a version of this problem that’s about dashboards, and it’s the boring version. Somebody built the wrong report; build a better one. The version Timo found in the first five minutes of his brainstorm is different, and I didn’t see it coming. Nothing here is technically broken. The tracking works. The data exists. Three competent people produce exactly what was asked for, every two weeks, on time. And the output is a decision made on instinct anyway. What’s actually broken, in his words, is that “mentally they’re in a bad place”— and that no fix laid on top of that frustration will survive it. That reframe changed the whole session. What follows is where we got to. The problem, as submitted Category: Data Analysis & Reporting. Submitted by: Anonymous. Context: B2C wellness app, Series A, ~250,000 monthly users. Three-person data team. Issue — Every two weeks, before sprint planning, the PM asks the data team what to prioritize. She receives a 40+ chart report from Mixpanel: session duration, feature adoption, retention cohorts, everything. She still can’t make a decision. “The data shows everything but recommends nothing.” She makes gut calls anyway, which feels wrong given how much data is on the table. The cost, per cycle: three people, one full day, to build it. Two to three days for her to read it. Then the deadline hits and the decision goes to whatever the CEO mentioned last week or whatever’s loudest in the support tickets. Trigger — Every sprint cycle. It came to a head last quarter: they redefined what the dashboard needed to show, the data team took seven weeks to rebuild the report, and she still couldn’t decide fast enough. Result: a wasted A/B run and a feature release that caused churn. (Additional context I happened to have, knowing the PM: the Mixpanel API connection was fine, but the aggregation kept collapsing under the event volume — a lot of manual correction every cycle, and eventually a need for different tooling.) Boundaries — Two-week sprint cycle, velocity must hold. Mixpanel stays. And the fix has to start working next sprint, not after a months-long transformation. Tension — Two sentences, both worth sitting with: “The data team is undervalued.” “We are not data informed, we are data paralyzed.” Clarity statement (where we landed live): Reverse-engineer from the decision backwards. Define what success of the product actually looks like, then design the smallest set of metrics that tells the PM where her biggest lever is before a sprint — rather than redesigning the report again. Our guest timo dechau 🕹🛠 — solo consultant, product and growth analytics. Based in Aalborg, Denmark (not Copenhagen, as he’d like on the record). Timo started in product and never really left it — most of what he does in data is still built on product principles. His path ran through classic tracking implementation, then an equally long stretch on the data warehouse side, and in the last two or three years into strategy, which he says he’d written off as “something for old people” until he found out what it’s actually for: setting expectations early enough to prevent the problems that show up later. He works with companies trying to get their warehouse stack into a shape that supports real product analytics, growth prediction, and marketing attribution, and he’s building a product that sits on top of the usual suspects — Amplitude and Mixpanel — to add a more strategic layer to product analytics. Connect with Timo: * Website: timodechau.com * LinkedIn: He posts two or three times a week, and describes LinkedIn as his first sounding board for ideas before they become blog posts or videos Where the two approaches met Timo’s first move: clear the room before you fix anything He got stuck on this for the longest part of his twenty minutes, and it’s the most useful thing in the episode. “I think the biggest problem is — they’re in a bad place mentally. Not business-wise, not from a data perspective. They’re collecting data, stuff is there. But mentally they’re in a bad place.” Frustration like this doesn’t start two months ago. It accumulates, on both sides, and by the time somebody writes “we are data paralyzed” into a problem submission, it has an audience — the product team has told other teams, management has heard about it, and the question in the room has quietly become why hasn’t this been solved already, when everyone else has solved it? (Very few companies have. Plenty claim to. That gap makes the environment harder, not easier.) So his quick win is a stop, not a start: Produce no sprint reporting at all for three sprints. Use the reclaimed capacity — a full day per cycle, per person — to build the replacement. And this cannot be agreed between the product and data teams alone. It goes up to management, it gets stated openly, and something gets delivered at the end. Otherwise the same trap closes again. Then: borrow the product strategy to narrow the scope Product is measurable in a thousand directions, which is why forty charts happened in the first place. Timo’s scope-narrowing device is the strategy nobody in data usually reads. There’s a business strategy. Product derives its strategy from it. Spend time with management and product understanding what they actually want to move in the next twelve months — then translate that movement into metrics. The payoff is political as much as analytical: “When we can come up with some metrics that show strategic progress, everyone is happy — even when we haven’t solved the core problem yet.” It takes pressure out of the room, it tells management whether their strategy is working, and it buys the credit needed to fix the rest properly. Then: measure outcomes, not interactions “One of the big issues of almost all product analytics setups is that it’s focusing on interactions and it’s not focusing on outcomes. Interaction is easier to track — there’s a button, we can click it, we can track it. But the real value comes when you take a step back and say: what are the outcomes of our wellness app? What do we want people to achieve?” Practically, that means event storming sessions to map the journey and identify value moments — even if it’s been done before — and then a shift in how results get presented: * Away from retention curves and cohort reports. Beautiful assets, genuinely useful, and readable by professional analysts. Not by a PM at 9am before sprint planning. * Towards metrics: a one-month, three-month, six-month retention rate. Put them on a time series. Cohort them later if you want. A PM can look at three of those and answer “did the last sprint move anything?” in about ten seconds. * Or user-state measurement, growth-model style: new → activated → active → at risk → dormant, and measure the movement between states. Activation rate, at-risk rate, and so on. Two constraints he names honestly. First, Mixpanel is a poor fit for this: “Mixpanel is not a metric-based tool. It’s an event data exploration tool” — metrics arrive late and aren’t first-class. Second, when the conversation stalls on “our tracking is bad,” the way out is often to stop tracking and instead derive from the product database. That’s source data. “Tracking cannot be good, because it happens in a browser.” Lior’s move: three questions every KPI has to survive I came at it from the other end — the architecture, not the tooling. Start from one to three KPIs that measure business value. Then put each one through three questions before it goes anywhere near a dashboard: * What action will this require of us? If a number moves and nobody does anything differently, it isn’t a KPI. It’s decoration. * Why do I need it? The purpose, stated. Day-7 retention exists so I can tell whether a feature is sticky enough to bring people back. Once that’s written down, question one has an answer. * Who owns the data, and who owns the KPI? Two different jobs. One is accuracy and lineage. The other is the definition, the calculation when it changes, and the communication. Then, in order: design the visualization and the filters → map back to the data sources so lineage is documented → communicate to management and get their buy-in, so nobody is surprised when the number moves → and test the whole thing with a simple CSV before asking anyone to build it. That last step matters more here than usual. The data team is already frustrated. Validating the concept by hand, before commissioning work, is the difference between one build and four. On ownership — my honest answer to Timo’s question about whether anyone ever agrees to own a KPI is a bad, mostly. But the principle holds: the requester is the domain expert, and the domain expert owns it. Handing it to the data team makes no sense, because as Timo put it, “they cannot make the smell test.” They’ll present a number, someone in the domain will say “that can’t be right,” and they’ll have no way to know who’s correct. And the decision book — an idea I took from Philip back in episode three and have since used myself. Alongside the metrics, write down what you do when each one moves. High level, not exhaustive. Each time a new scenario shows up, add it. It’s what turns “the number went down” into a sprint conversation instead of an argument. Three insights 1. Minimal, strategy-linked metrics get you halfway on their own. Both of us arrived at a small number — Timo from the strategy end, me from the KPI end. The convergence matters less than the discipline: don’t have more metrics than you

    Forty Charts, No Decision
  2. Sep 2

    Four Dashboards, Four Different Forecasts

    Data Breakthroughs — S2E1: Four Dashboards, Four Different Forecasts Real-world data problem solving in action. Guest Koray and host Lior Barak open a community-submitted challenge for the first time on the recording — no prep, no script. Problem category: Business Intelligence & Dashboarding Runtime: [[34]] minutes The challenge: A Series B B2B SaaS company has four dashboards reporting the current-quarter sales pipeline, each with a different total — and after a public disagreement between the CFO and the sales manager, the VP has stopped trusting all of them. The solution: Separate the data problem from the forecasting-logic problem by re-running all four dashboards against a closed historical quarter, then rebuild forward from one agreed calculation with the business in the room at every step. Key takeaways: Trust is the real damage. It goes fast and comes back slowly. Test against a closed quarter before touching anything — it tells you whether the split is in the data or in the assumptions. Different teams can keep different operational definitions. They cannot keep unagreed ones. Disclaimer: This podcast is for inspiration and educational purposes. Solutions discussed are general approaches — adapt them to your context and constraints. Music: "Calisson" courtesy of Riverside This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit impactoperations.substack.com

    Four Dashboards, Four Different Forecasts
  3. 12/17/2025

    Why Simple Data Integrations Take Months & How to Fix Them (feat. Ilya Vladimirskiy)

    Episode Summary I’m incredibly excited to have Ilya back for the Season 1 finale. He was my very first guest in the pilot episode, and honestly, he helped me figure out this format - what works, what doesn’t, how to make the collaborative problem-solving feel authentic. So it felt only right to close this season by bringing him back full circle. In this finale, we tackle a frustratingly common scenario: a marketing stakeholder needs data from a new platform for critical quarterly forecasting, but the data team estimates four months to build the connector. Meanwhile, two hours disappear every morning into manual CSV downloads, cleanups, and copy-paste operations, a process that’s already caused a 50% error in pipeline analysis. What seems like a technical integration problem quickly reveals itself as something much deeper: an organizational breakdown in ownership, communication, and mutual understanding between business stakeholders and data teams. And true to form, Ilya immediately zeroes in on the people side of things. Problem Category: Data Integration & ETLRuntime: 40 minutes The Problem Submitted by: Anonymous Marketing ProfessionalIndustry Context: Company with established data infrastructure and quarterly business planning cycles Problem Framework Issue: Need data from the new marketing automation platform for quarterly forecasting, but building a connector will take four months, according to the data team. Currently manually downloading CSV files daily. Trigger: Quarterly planning cycle starts in 8 weeks. Currently spending 120 minutes every morning downloading, cleaning, and manually importing marketing data files. Last week, a copy-paste error threw off pipeline analysis by 50% - only caught because the numbers seemed unrealistic. Tension: The data team focuses on building robust, enterprise-grade connectors that take months to develop properly. While understanding their approach, there’s an immediate business need that can’t wait for the perfect solution. The manual process is unsustainable and risky, but the data is critical for business planning. Boundaries: * Cannot change the quarterly planning timeline (set by business cycle) * The marketing platform was selected by leadership and cannot be changed * The data team has limited capacity and other priorities * A budget exists for reasonable interim solutions * Must maintain data quality standards for forecasting accuracy Tech Stack: New marketing automation platform with CSV export capability, central data warehouse for forecasting (specific tools not disclosed) Clarity Statement: Need an interim solution to get marketing platform data into the data warehouse reliably within the next 8 weeks, without waiting for the full enterprise connector that will take 4 months. Our Guest IlyaFractional Head of Data & Data Leadership Consultant Ilya brings over 15 years of data experience, having led data functions at companies like Ada Health (symptom checker app), and various startups and scale-ups across Berlin and Munich. Originally from Moscow with a background in computational mathematics, he moved to Germany in 2002 and transitioned from database research to hands-on data engineering and leadership roles. After the biotech winter impacted Ada Health, Ilya pivoted to fractional and interim data leadership, helping companies build data platforms and teams across different domains and stages. Special Note: Ilya was our very first guest in the pilot episode and returns to close out Season 1, bringing his people-first philosophy full circle. Connect with Ilya: * LinkedIn: https://www.linkedin.com/in/bkmy43/ * YouTube: https://www.youtube.com/@lab4.berlin (Data leadership discussions while smoking pipes - yes, really, and it’s worth checking out) * Website: https://www.lab4.berlin/ The Breakthrough Discussion Initial Reactions Ilya and I immediately recognized this as a people problem disguised as a technical problem. As Ilya put it: “From most failing projects and situations like this, I rarely saw the root cause was technology or tools.” The four-month estimate raised red flags for both of us. As Ilya observed, connecting to an API of an existing marketing tool shouldn’t take four months - what’s likely happening is that “building a connector” actually means the entire pipeline: getting the data, integrating it into the company data model, and delivering it through BI tools with proper business logic. The Real Problem Through the reflection break and collaborative discussion, we identified the core issues: * Communication Breakdown: The data team likely down-prioritized this request because other initiatives have a higher business impact, but they haven’t articulated this clearly. The stakeholder hears “four months” without understanding what’s blocking it. * Ownership Confusion: It’s unclear who owns the decision about prioritization and tradeoffs. Without clear ownership, every request becomes a negotiation rather than a strategic decision. * Missing Context: The data team probably doesn’t understand the business impact of the delay (corrupted forecasts, wasted ad spend, strategic planning delays). The stakeholder doesn’t understand what the data team is actually building and why it takes time. * Us vs Them Dynamic: The situation has devolved into adversarial positioning - “the data team won’t help me” versus “stakeholders want everything immediately” - rather than collaborative problem-solving. The Solution Approach Rather than a single technical fix, our discussion produced a multi-layered solution: Immediate Relief (Week 1-2): Ask one of the engineers to build a simple Python script or similar automation. This doesn’t need to be production-grade infrastructure - just something that reliably pulls the CSV, does basic transformation, and loads it into the warehouse. This can be a 2-3 day task rather than a 4-month project. Transparency & Context (Week 2-3): Create a visible initiative backlog overview showing everything the data team is working on. When someone says “it will take four months,” they should be able to show exactly what’s blocking it and why those other priorities matter more. This isn’t about justifying delays - it’s about enabling informed decisions. Rational Decision Framework (Week 3-4): Develop a structured way to articulate both the cost of building solutions and the business impact of delays. Put numbers on the table: What does two hours of manual work daily cost? What’s the risk value of potential forecast errors? What’s the opportunity cost of the data team working on this versus their current priorities? Strategic Alignment (Ongoing): Establish clear ownership and a prioritization process that considers both technical complexity and business impact. This isn’t about the data team gatekeeping or stakeholders demanding - it’s about having a framework where tradeoffs are visible and decisions are rational. Key Takeaways 3 Critical Insights * This is an Organizational Problem, Not a Technical One: The four-month timeline isn’t about technical complexity - it’s about priorities, communication, and organizational dynamics. The actual technical work of connecting to an API could be done much faster if approached as a quick automation rather than an enterprise-grade infrastructure. * The “Us vs Them” Dynamic Is Killing Efficiency: When stakeholders and data teams position themselves as adversaries rather than collaborators, every interaction becomes a negotiation. The marketing person sees the data team as obstructionist; the data team sees stakeholders as demanding and unrealistic. Neither side wins in this dynamic, and the business suffers. * Ownership Clarity Is Essential: Without clear ownership of prioritization decisions, every data request becomes contested territory. Someone needs to own the decision about whether a four-month wait is acceptable given the business impact, and that person needs visibility into both the technical constraints and business consequences. 4 Action Items For the Problem Submitter (and anyone in similar situations): * Request a Quick Automation Script (This Week) - Ask a data engineer to build a simple Python script or similar automation that pulls the CSV, does basic transformation, and loads it into your warehouse. Make it clear this doesn’t need to be production-grade infrastructure - just something reliable enough to bridge the gap. Timeline: 2-3 days of engineering time. * Create Initiative Backlog Visibility (Week 2) - Work with the data team to create a visible overview of all current initiatives. When told something will take four months, you should understand what’s blocking it and why those priorities were chosen. This isn’t about challenging their decisions - it’s about having context for informed discussion. * Articulate Cost and Impact With Numbers (Week 3) - Document the actual business impact: two hours daily of manual work (cost it out by salary), risk of forecast errors (quantify the potential impact), strategic planning delays (what decisions are being made without this data?). Similarly, ask the data team to articulate what they’d need to deprioritize to tackle this sooner. * Establish Ongoing Prioritization Framework (Week 4+) - Work with leadership to create a clear process for prioritizing data work that considers both technical complexity and business impact. Identify who owns these decisions and ensure they have visibility into both technical constraints and business consequences. This prevents future “four months” surprises. Episode Highlights * 02:00 - Problem reveal: Four months for a marketing platform connector * 06:30 - Ilya’s immediate diagnosis: “This is a people problem, not a technical one” * 14:45 - Post-reflection discussion: Unpacking the communication breakdown * 24:30 - The ownership question: Who actually decides priorities? * 31:00 - Quick wins vs long-term solutions

    Why Simple Data Integrations Take Months & How to Fix Them (feat. Ilya Vladimirskiy)
  4. 12/10/2025

    How to Bridge the Data-Experience Gap & Gain Executive Buy-In (feat. Tiankai Feng)

    Data Breakthroughs - Episode 10: When Data Meets Decades of Experience Real-world data problem solving in action! Tiankai Feng (Director of Data & AI Strategy at ThoughtWorks, author of "Humanizing Data Strategy" and "Humanizing AI Strategy") and host Lior Barak tackle a manufacturing company where plant managers with 20+ years of experience resist a modern data platform. Problem Category: Organizational Data Strategy / Change ManagementRuntime: 32 minutes The Challenge: A family-owned manufacturer invested heavily in data infrastructure, but plant managers still make decisions based on "what happened last time" and gut instincts, creating a divide between analytics teams and operations. The Solution: Transform data from replacement threat to support tool through co-creation, clear communication about expertise-data synergy, and defining decision-making rules with proper incentives. Key Takeaways: Expertise vs. data is always a tension field - communicate how they work hand-in-hand, not against each other People don't use things they didn't help create - co-creation is essential for adoption Success metrics must reflect both short-term and long-term value to align incentives properly Guest: Tiankai Feng, Director of Data & AI Strategy at ThoughtWorks Author of "Humanizing Data Strategy" and "Humanizing AI Strategy" Connect: https://www.linkedin.com/in/tiankaifeng/ Get Involved: Submit your data problem or Become a guest: https://data-breakthroughs-podcast.cookingdata.blog/ Join the conversation: #DataBreakthrough Full show notes: https://data-breakthroughs-podcast.cookingdata.blog/ Disclaimer: This podcast is for inspiration and educational purposes. Solutions discussed are general approaches - adapt them to your specific context and constraints. Music: "Calisson" courtesy of Riverside This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit impactoperations.substack.com

  5. 11/26/2025

    How to Manage Analytics Overload & Build Effective Dashboards (feat. Eva Schreyer)

    Data Breakthroughs - Episode 09: When Analytics Becomes a Dashboard Factory Real-world data problem solving in action! Eva Schreyer (Head of Data & Analytics at Neugelb/Commerzbank) and host Lior Barak tackle a community-submitted challenge about analytics overload for the first time during recording. Problem Category: Business Intelligence & Dashboarding / Organizational Data StrategyRuntime: 37 minutes The Challenge: A product team drowns in 40+ charts per report while struggling to make data-driven decisions, creating a disconnect between analytics investment and business value. The Solution: Transform from dashboard factory to strategic partner through executive alignment, monetizing requests, and prioritizing deep-dive analyses over generic reporting. Key Takeaways: Too much data doesn't mean good decisions-relevance matters more than volume Making stakeholders understand the cost of requests (in time/effort) dramatically improves prioritization Ask "what will you do differently when this metric changes?" to identify truly actionable insights Guest: Eva Schreyer, Head of Data & Analytics at Neugelb (Commerzbank) Connect: https://www.linkedin.com/in/eva-schreyer/ Get Involved: Submit your data problem: https://data-breakthroughs-podcast.cookingdata.blog/ Become a guest: https://data-breakthroughs-podcast.cookingdata.blog/ Join the conversation: #DataBreakthrough Full show notes & visual diagrams: https://data-breakthroughs-podcast.cookingdata.blog/ Disclaimer: This podcast is for inspiration and educational purposes. Solutions discussed are general approaches - adapt them to your specific context and constraints. Music: "Calisson" courtesy of Riverside This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit impactoperations.substack.com

  6. 11/12/2025

    Why Your Perfect Model Fails in Production: The Accuracy Paradox (feat. Irena Bojarovska)

    Data Breakthroughs - Episode 8: The Office Kitchen Paradox Real-world data problem solving in action! Irena Bojarovska and host Lior Barak tackle a community-submitted challenge for the first time during the recording. Problem Category: Machine Learning & AI Implementation Runtime: 50 minutes The Challenge: A hackathon team built a smart kitchen demand forecasting model with 91% accuracy, but the company is still throwing away 20-25% of fresh products weekly while running out of popular items. The Solution: The breakthrough isn't about fixing the model, it's about fixing the data. The model is missing critical inputs (office attendance, special events) and is operating blindly due to data quality problems. The real solution combines better data, human-AI collaboration, and proper A/B testing. Key Takeaways: • Model accuracy ≠ real-world performance (91% test accuracy doesn't guarantee waste reduction) • Data quality and contextual information are your foundation (garbage in, garbage out) • Humans should augment the model, not be replaced by it (hybrid approach wins) Guest: Irena Bojarovska, Data Scientist at Zalando SEConnect: https://www.linkedin.com/in/irenabojarovska/ Get Involved: Submit your data problem: https://data-breakthroughs-podcast.cookingdata.blog/submit-problem Become a guest: https://data-breakthroughs-podcast.cookingdata.blog/become-guest Join the conversation: #DataBreakthrough Full show notes & visual diagrams: https://wabi-sabi-data-newsletter.com [or your actual newsletter link] Figma Board: https://www.figma.com/board/jfC4ipNvd8zSPIyZreEten/Irena-Bojarovska?node-id=1-14&t=Q46O2Ae9yuRHZRwy-1 Disclaimer: This podcast is for inspiration and educational purposes. Solutions discussed are general approaches; adapt them to your specific context and constraints. Music: "Calisson" courtesy of Riverside This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit impactoperations.substack.com

  7. 11/12/2025

    How to Compete Against AI-Powered Competitors With Limited Resources (feat. Jon Cooke)

    Data Breakthroughs - Episode 7: Small Company vs AI Giants Real-world data problem solving in action! Jon Cooke (Founder of Dataception) and host Lior Barak tackle a classic David vs. Goliath scenario for the first time during the recording. Problem Category: Data Strategy & Customer AnalyticsRuntime: 36 minutes The Challenge: Small German seed company (7 people) with 600+ product varieties, 4 years of customer data, and 30 years of gardening expertise. They're losing to giants who use algorithms for personalized recommendations. Conversion rate: 2.1%. Sent tomato seeds in December while competitors suggested microgreens and winter planning guides. They have incredible data and domain knowledge - but no idea how to compete with automated personalization. The Solution: You don't need massive tech teams. Start with customer segmentation workshops, map buying journeys, understand your data quality, build a simple recommendation engine (could be done in half a day), and test with friendly customers. The institutional knowledge trapped in people's heads is your competitive advantage - you just need to capture and automate it. Key Takeaways: Understand customers and segments first - technology second Data quality dictates approach: good data = ML models, poor data = heuristic rules Simple models beat no models - you don't need world-class data scientists This is a business process problem with AI tools, not an AI problem Small teams can compete by moving fast and testing with customers Guest: Jon Cooke, Founder of Dataception20 years in data & AI | Former Databricks Solutions Architecture Lead | Ex-PwCExpert in data products, GenAI, and knowledge graphs Connect with Jon: Website: https://dataception.com LinkedIn Company: Dataception Get Involved: Submit your data problem: https://data-breakthroughs-podcast.cookingdata.blog/submit-problemBecome a guest: https://data-breakthroughs-podcast.cookingdata.blog/become-guestJoin the conversation: #DataBreakthrough Full show notes & visual diagrams: [Link to newsletter version] Disclaimer: This podcast is for inspiration and educational purposes. Solutions discussed are general frameworks - adapt them to your specific context and constraints. Music: "Calisson" courtesy of Riverside This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit impactoperations.substack.com

  8. 10/29/2025

    How to Build Explainable AI Models & Prevent Regulatory Disasters (feat. Elizabeth Press)

    Data Breakthroughs - Episode 6: AI Bias & Explainability Crisis Real-world data problem solving in action! Elizabeth Press (Founder of D3M Labs, Deputy Chief Digital Officer at CHESCO) and host Lior Barak tackle one of AI’s most critical challenges for the first time during the recording. Problem Category: Machine Learning & AI Implementation / AI EthicsRuntime: 36 minutes The Challenge: A deep learning loan approval model improved accuracy by 23% and reduced processing time from days to minutes. Business results? Phenomenal. The problem? It systematically denies qualified applicants in certain zip codes at 40% higher rates - and the model is a black box that can’t explain individual decisions. Regulatory examination in 12 weeks. Potential discrimination lawsuits are looming. The Solution: Not all use cases should use unexplainable AI. Return to statistical fundamentals (logistic regression), implement hybrid human-in-loop systems, create cross-functional teams involving legal from day one, build test boxes for domain validation, and establish decision logs. Sometimes boring statistics beat sexy deep learning. Key Takeaways: High-stakes decisions (loans, justice, healthcare) should never use unexplainable black box models Involve legal, PR, and domain experts from the start - not retroactively Bias is quantifiable through business metrics (churn, customer complaints, defaults) You must be able to explain your model - without it, you run into catastrophic risks Speed means nothing if accuracy and ethics are compromised Guest: Elizabeth Press, Founder of D3M Labs | Deputy Chief Digital Officer at CHESCO. Former data leader | Taught “Profitable AI” at Hasso Plattner InstituteBackground in financial risk management and credit rating models Connect with Elizabeth: D3M Labs: https://www.linkedin.com/company/d3m-associates/posts/?feedView=all YouTube: D3M Labs channel LinkedIn: Elizabeth’s profile Focus: Profitable and secure digital business Get Involved: Submit your data problem: https://data-breakthroughs-podcast.cookingdata.blog/submit-problemBecome a guest: https://data-breakthroughs-podcast.cookingdata.blog/become-guestJoin the conversation: #DataBreakthrough Full show notes & visual diagrams: [Link to newsletter version] Disclaimer: This podcast is for educational and inspirational purposes. Neither host nor guest is/are lawyer. AI ethics and legal compliance require professional legal counsel. Solutions discussed are general frameworks - adapt them to your specific context, regulations, and legal requirements. Music: “Calisson” courtesy of Riverside This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit impactoperations.substack.com

Trailers

About

A podcast where data experts solve real-world operational challenges submitted by listeners. Each episode tackles a fresh problem, delivering actionable solutions, key insights, and implementation steps to help data professionals overcome barriers and create business value. impactoperations.substack.com