Data in Biotech

CorrDyn

Data in Biotech is a fortnightly podcast exploring how companies leverage data to drive innovation in life sciences.  Every two weeks, Ross Katz, Principal and Data Science Lead at CorrDyn, sits down with an expert from the world of biotechnology to understand how they use data science to solve technical challenges, streamline operations, and further innovation in their business.  You can learn more about CorrDyn - an enterprise data specialist that enables excellent companies to make smarter strategic decisions - at www.corrdyn.com

  1. 3 hr ago

    Why Biotech Talks About AI But Won't Pay for the Data It Needs

    Everyone in biotech agrees AI needs more data. Almost no one is willing to pay for it. If you're trying to build or buy a biotech AI model, you've hit the same wall: predictive performance depends on data your budget doesn't cover, and nobody in the field seems willing to close that gap. John Androsavich runs Ginkgo Datapoints, the bio AI data arm of Ginkgo Bioworks. He trained as an RNA scientist, spent years on the pharma side deciding which technologies were worth buying, and now sells the raw biological data everyone claims to want. Ross and John get into why biotech spends a fraction of what tech spends on data, how automation dropped ADME testing to $199 a compound, and what that unlocks for drug discovery pipelines and data science in biotech more broadly. You'll hear why single-cell foundation models don't scale the way the field expected, and how GPT-5 designed its own lab experiments inside an autonomous facility. This one's for data and analytics leaders in biotech who need a clearer read on where to spend on data generation, and where the field is still guessing. It's less useful if you're after a general AI overview with no biotech specifics. Key Takeaways - One Meta investment in a data-labelling vendor outweighs a full year of AI drug discovery venture funding combined, and dwarfs the entire single-cell data market. Biotech's data spend looks nothing like tech's. - Ginkgo's ADME-1 offering runs at roughly a tenth of standard pricing, which is changing when and how much companies test. Teams are now running full tier-one panels earlier instead of triaging molecules before they've generated the negative data models need. - A recent Microsoft Research paper found single-cell foundation model learning saturates at 200,000 to 2 million cells, out of a possible 20 million. Volume alone isn't the lever people assumed it was. - GPT-5 wrote its own experimental protocols for optimising cell-free protein expression, ran them through Ginkgo's autonomous Nebula lab, and hit the lowest price-per-titer ever recorded in the field. Chapter Markers 00:00 Introducing John Androsavich and Ginkgo Datapoints 01:12 Why Ginkgo launched a bio AI data business 05:03 Which companies benefit most from Datapoints 06:31 The paradox: everyone wants data, no one pays 09:00 How automation drives ADME-1's $199 price point 12:59 Testing the Jevons paradox in biotech data buying 16:05 Do we actually know biotech AI's scaling laws? 20:54 Why foundation model builders resist more data 24:59 What an empirical bake-off for bio AI could look like 29:32 The case against sitting on the sidelines 33:26 Inside the Virtual Cell Pharmacology Initiative 41:57 Where VCP fits among other virtual cell projects 44:50 The Antibody Developability Consortium with Apheris 53:57 Autonomous labs and GPT-5 designing its own experiments 59:38 Advice for mid-stage biotech data strategy 01:01:31 Final thoughts on where bio AI investment is heading Useful Links & Resources - Ginkgo Bioworks: [ginkgobioworks.com](https://www.ginkgobioworks.com) - Related episode: Apheris CEO Robin Rohm on federated co-folding (Data in Biotech) - Related episode: Eliza Appel on Lilly's TuneLab and federated learning (Data in Biotech) - CorrDyn: [corrdyn.com](https://www.corrdyn.com) Connect With the Show - Host LinkedIn (Ross Katz): [linkedin.com/in/b-ross-katz](https://www.linkedin.com/in/b-ross-katz/) - Host X: [x.com/brosskatz](https://x.com/brosskatz) - CorrDyn LinkedIn: [linkedin.com/company/corrdyn](https://www.linkedin.com/company/corrdyn/) Where does your organisation sit on the data investment paralysis John describes? Are you waiting for someone else to prove the scaling laws first, or are you buying the data now? Drop your take in the comments. Visit corrdyn.com to learn how CorrDyn can help your organisation extract value from data. #DataInBiotech #BiotechAI #DrugDiscovery #DataScience #GinkgoBioworks

  2. 20 Jul

    Beyond Language: Why Drug Discovery Needs Physical AI, Not Just Large Language Models

    In this episode of Data in Biotech, host Ross Katz sits down with Woody Sherman, Founder and Chief Innovation Officer at PsiThera, for a conversation on why AI can transform drug discovery's paperwork and code while barely touching the hardest part of the problem: the molecules themselves. Woody's career runs through physical chemistry at MIT; over a decade at Schrödinger building tools the industry still relies on; founding Silicon Therapeutics (where his team took a small molecule STING agonist from concept to clinic in roughly three years); scaling that platform after Roivant's acquisition; and now leading PsiThera's effort to build oral small molecules for immunology targets that today are only reachable with injectable biologics. The conversation digs into why large language models excel at automation, coding, and regulatory writing but hit a wall when the task is predicting how a molecule behaves, what "physical AI" actually means as a category distinct from both LLMs and traditional physics-based simulation, and why representing molecules as quantum mechanical objects rather than text strings or 2D graphs changes what's predictable.  Woody also walks through the STING program in detail, why the field's excitement over fast co-folding models like Boltz needs a strong dose of skepticism, and what it takes to build a database and team culture where chemists, biologists, and data scientists can actually understand each other. What you'll learn in this episode: >> Why the contradiction of "AI is transforming drug discovery" and "drugs still take a decade and billions of dollars" can both be true at once. >> How Silicon Therapeutics engineered a small molecule STING agonist to dimerize itself through a quantum mechanical interaction that had never been designed for before. >> What "physical AI" means as a new category built on embeddings from orbital-level, quantum mechanical representations of molecules, rather than language tokens or force-field simulations. >> Why molecular representation is the whole game: the limitations of SMILES strings and 2D graphs versus true 3D, quantum mechanical embeddings like PsiThera's Psiformer model >> Why a widely publicized claim of near-FEP-quality binding affinity at 1,000x the speed didn't hold up under scrutiny. >> How PsiThera captures not just simulation and wet lab data but human chemist judgment and reasoning as structured data, and why building a shared vocabulary across computational and experimental teams is as important as any model. Meet our guest: Woody Sherman, PhD, is Founder and Chief Innovation Officer at PsiThera, a biotechnology company designing oral small molecule drugs for immunology and inflammatory diseases, starting with the TNF superfamily. His career spans physical chemistry research at MIT, more than a decade at Schrödinger developing computational drug discovery tools, founding Silicon Therapeutics (acquired by Roivant), and leading the platform's evolution through PsiThera today. He has published more than 100 peer-reviewed papers spanning molecular dynamics, quantum mechanics, free energy simulations, and machine learning for drug design. Connect with Woody Sherman on LinkedIn: https://www.linkedin.com/in/woodysherman/ About the host: Ross Katz is Principal and Data Science Lead at CorrDyn. Ross specializes in building intelligent data systems that empower biotech and healthcare organizations to extract insights and drive innovation. Connect with Ross Katz on LinkedIn: https://www.linkedin.com/in/b-ross-katz/ Sponsored by… This episode is brought to you by CorrDyn, the leader in data-driven solutions for biotech and healthcare. Discover how CorrDyn is helping organizations turn data into breakthroughs at CorrDyn. https://www.linkedin.com/company/corrdyn/

  3. 17 Jun

    Synthesizable by Design: Rethinking AI's Role in Small Molecule Drug Discovery

    In this episode of Data in Biotech, host Ross Katz sits down with Paul Finn, Chief Scientific Officer at Oxford Drug Design, for a conversation on what it actually takes to find a drug molecule that works not just on paper but also in the lab, in the cell, and, ultimately, in the clinic. Paul brings four decades of experience across what became GSK, Pfizer, and a series of Oxford-area spinouts and has shepherded a compound all the way to a marketed drug. That perspective gives him a particular kind of skepticism toward AI results that look too good to be true because he's done the work of checking whether they are. The conversation moves through synthesizability as a first-class constraint, why chemistry has proven so much harder for AI than biology, how 3D molecular representation gets closer to the physics that actually matters, and what rigorous multi-parameter optimization looks like when you're trying to kill cancer cells and drug-resistant bacteria at the same time. What you'll learn in this episode: >> Why synthesizability is chronically underestimated and why changing a single atom in a structure can take a molecule from trivially easy to make to practically impossible >> How Oxford Drug Design constrains the generative search to reaction schemes and purchasable building blocks, and why that chemical space is still so vast that novelty is not meaningfully sacrificed >> Why most generative AI models learn from a 2D string representation of a molecule; two steps removed from the 3D physics that govern how a drug actually binds to its target >> How Bayesian optimization over reagent space, rather than molecular space, allows an active learning loop to focus on the structural patterns associated with activity >> Why benchmarking complex models against simple ones is the discipline that exposes false correlations and why Paul and his co-authors were able to recover the Halicin result using methods decades older than deep learning >> What a pharma company should actually ask an AI drug discovery vendor before buying what they're selling Meet our guest: Paul Finn is Chief Scientific Officer at Oxford Drug Design, a computational drug discovery company with roots in Oxford's chemistry department. His career spans over 40 years of computational drug discovery, from early structure-activity modeling in the 1980s through to modern generative AI methods, with deep experience at what became GSK and Pfizer before moving into the Oxford spinout ecosystem. At Oxford Drug Design, Paul leads internal programs in oncology and antibacterial resistance, combining novel computational methods with a rigorous, synthesizability-first approach to multi-parameter optimization. Connect with Paul Finn on LinkedIn: https://uk.linkedin.com/in/paul-finn-2250616 About the host: Ross Katz is Principal and Data Science Lead at CorrDyn. Ross specializes in building intelligent data systems that empower biotech and healthcare organizations to extract insights and drive innovation. Connect with Ross Katz on LinkedIn: https://www.linkedin.com/in/b-ross-katz/ Connect with us: Follow the podcast for more insightful discussions on the latest in biotech and data science.Subscribe and leave a review if you enjoyed this episode! Sponsored by… This episode is brought to you by CorrDyn, the leader in data-driven solutions for biotech and healthcare. Discover how CorrDyn is helping organizations turn data into breakthroughs at CorrDyn. https://www.linkedin.com/company/corrdyn/

  4. 2 Jun

    From Tissue to Mechanism to Decision: Building AI for Computational Oncology

    In this episode of Data in Biotech, host Ross Katz sits down with Arvind Rao, Professor of Computational Medicine and Bioinformatics at the University of Michigan, for a discussion on the gap between what biomedical AI can do and what it can reliably be trusted to do in clinical practice. Arvind's research sits at the intersection of computational oncology and AI governance and his lab works across H&E histopathology, multiplex immunofluorescence, spatial transcriptomics, and single-cell RNA sequencing, not just to build predictive models, but to understand the full lifecycle from data to model to inference, and to ask where that lifecycle can be trusted and where it can't.  The conversation moves through two of his recent papers on SPIFEE, a graph-based framework that replaces scalar interaction scores in the tumor microenvironment with spatially resolved functional representations, and a multimodal framework that traces a path from stained tissue slides to nominated drug targets via morphological pattern discovery and spatial transcriptomic mapping.  What you’ll learn in this episode:  >> Why the field's central failure is not algorithmic but translational and the gap between a model that performs well on a benchmark and one that can be consistently trusted in a high-stakes clinical setting  >> How SPIFEE replaces the conventional scalar edge representation of cell-cell interactions in the tumor microenvironment with spatially resolved functional edges >> How Arvind's multimodal framework moves from H&E pathology slides labeled with clinical outcomes, through morphological pattern discovery via multiple instance learning, to spatial transcriptomic mapping, to the nomination of molecular mechanisms and actionable drug targets >> Why Goodhart's Law applies directly to foundation model evaluation in biology  >> What the AI literacy gap costs when it goes unaddressed in healthcare and pharma organizations  Meet our guest: Arvind Rao is a Professor of Computational Medicine and Bioinformatics, with a joint appointment in Radiation Oncology, at the University of Michigan. His research focuses on establishing trust in biomedical AI predictions across the full data-to-decision pipeline, integrating H&E histopathology, spatial transcriptomics, multiplex immunofluorescence, and single-cell RNA sequencing to build models that are predictive, interpretable, and biologically credible. Alongside his research, Arvind develops AI literacy programs for healthcare and pharma professionals, helping clinical and procurement teams evaluate and govern AI systems with the rigor those decisions demand. Connect with Arvind Rao on LinkedIn: https://www.linkedin.com/in/arvind-rao-3301301ba/ About the host: Ross Katz is Principal and Data Science Lead at CorrDyn. Ross specializes in building intelligent data systems that empower biotech and healthcare organizations to extract insights and drive innovation. Connect with Ross Katz on LinkedIn: https://www.linkedin.com/in/b-ross-katz/ Connect with us: Follow the podcast for more insightful discussions on the latest in biotech and data science.Subscribe and leave a review if you enjoyed this episode! Sponsored by… This episode is brought to you by CorrDyn, the leader in data-driven solutions for biotech and healthcare. Discover how CorrDyn is helping organizations turn data into breakthroughs at CorrDyn. https://www.linkedin.com/company/corrdyn/

  5. 13 May

    Cavities in the Data: Building FDA-Cleared AI for Dental Imaging with Overjet

    In this episode of Data in Biotech, host Ross Katz sits down with Sadegh Salehi, Director of Research and Principal Scientist at Overjet, to explore what rigorous model evaluation actually looks like when the stakes are clinical.  Overjet builds FDA-cleared vision models that detect and quantify dental disease across billions of X-ray images from thousands of practices - a data problem with a staggering number of dimensions. Thirty-two teeth per adult patient, each with different morphology. Multiple image types capturing different anatomy. Fifteen to twenty sensor manufacturers producing perceptually distinct images, each with different contrast, resolution, and noise characteristics. And disease severity distributions ranging from barely visible early-stage decay to obvious pathology.  Sadegh walks through what it takes to evaluate models responsibly across all of those dimensions and discusses why aggregate metrics like F1 score can mask catastrophic failures on specific subgroups, how models find and exploit shortcuts in training data, and why the same flawed sampling that creates gaps in your training set also creates them in your test set.  He also traces Overjet's architectural evolution from over twenty narrow task-specific models to a single foundation model they call Unity, explains how treatment plan procedure codes provide a noisy but real production feedback signal, and describes how Overjet became one of the first companies to secure the FDA's Predetermined Change Control Plan (a framework that allows model updates without filing a new clearance each time.) What you’ll learn in this episode:  >> Why aggregate evaluation metrics are insufficient for high-stakes medical AI  >> How models exploit shortcuts in training data: if all images from a rare sensor in the training set happen to be healthy, the model doesn't learn to read that sensor, it learns that the sensor means healthy, bypassing the visual task entirely and producing systematic false negatives in production >> How Overjet evolved from over twenty narrow, sensor-specific and indication-specific models into a single foundation model called Unity, using noisy labels generated by the small models as the training signal for a much larger backbone, then building independent prediction heads for each clinical indication on top of it >> Why the decision to keep prediction heads architecturally independent from one another was driven as much by FDA regulatory strategy as by modeling considerations >> How Overjet uses dental treatment plan procedure codes as a production monitoring signal Meet our guest: Sadegh Salehi is Director of Research and Principal Scientist at Overjet, where he leads the team responsible for building, evaluating, and deploying FDA-cleared vision models for dental disease detection and quantification.  Connect with Sadegh Salehi on LinkedIn: https://www.linkedin.com/in/sadegh-salehi/ About the host: Ross Katz is Principal and Data Science Lead at CorrDyn. Ross specializes in building intelligent data systems that empower biotech and healthcare organizations to extract insights and drive innovation. Connect with Ross Katz on LinkedIn: https://www.linkedin.com/in/b-ross-katz/ Connect with us: Follow the podcast for more insightful discussions on the latest in biotech and data science.Subscribe and leave a review if you enjoyed this episode! Sponsored by… This episode is brought to you by CorrDyn, the leader in data-driven solutions for biotech and healthcare. Discover how CorrDyn is helping organizations turn data into breakthroughs at CorrDyn. https://www.linkedin.com/company/corrdyn/

  6. 30 Apr

    Data as a Moat: Why Biotech's Most Valuable Asset is Buried in a Hard Drive

    In this episode of Data in Biotech, host Ross Katz sits down with Jesse Johnson, founder of Merelogic, a software consulting firm specializing in data infrastructure for biotech organizations.  Jesse brings a rare perspective to the conversation: having built data systems at Google where engineers control the data collection function end to end, before moving into biotech, where the biology does what it wants and bench scientists, not engineers, generate the data.  The result is a grounded, pragmatic take on one of the most consequential and underappreciated questions in life sciences right now: as bio foundation models fundamentally change the value equation for experimental data, are biotech labs structured to capture that value?  Jesse argues the answer is usually no and that the fix is less technical than most assume. It doesn't require a production-grade data pipeline or a cloud architecture. It requires lightweight, human-readable standard operating procedures, clear expectations between computational and wet lab teams, and a data strategy designed not just for the questions you're asking today, but for the ones you don't yet know you'll need to ask. What you’ll learn in this episode:  >> Why the transition from tech to biotech requires a fundamental reset of assumptions about data infrastructure and why the biggest difference isn't technical, it's organizational. >> How bio foundation models have flipped the value equation for experimental data by reducing the cost of organizing it while dramatically increasing the potential return >> How the strategic value of proprietary data is evolving in the biotech ecosystem, from Tahoe Therapeutics building an acquirable single-cell dataset to Eli Lilly's Lowe lab using data as currency for partnerships  >> Why electronic lab notebooks aren't going anywhere and how the real question facing biotech software teams isn't whether to use an ELN, but how to balance schema rigidity against the flexibility required for the long tail of one-off exploratory assays that no automation pipeline will ever fully capture Meet our guest: Jesse Johnson is the founder of Merelogic, a software consulting firm that works with biotech and biopharma organizations on data infrastructure and data operations strategy. Jesse writes regularly about data strategy for biotech on his Substack, covering topics from bio foundation model adoption to the evolving role of electronic lab notebooks in an AI-augmented research environment. Connect with Jesse Johnson on LinkedIn: https://www.linkedin.com/in/jesse-johnson-biotech/ Follow Merelogic on Linkedin: https://www.linkedin.com/company/merelogic/ About the host: Ross Katz is Principal and Data Science Lead at CorrDyn. Ross specializes in building intelligent data systems that empower biotech and healthcare organizations to extract insights and drive innovation. Connect with Ross Katz on LinkedIn: https://www.linkedin.com/in/b-ross-katz/ Connect with us: Follow the podcast for more insightful discussions on the latest in biotech and data science.Subscribe and leave a review if you enjoyed this episode! Sponsored by… This episode is brought to you by CorrDyn, the leader in data-driven solutions for biotech and healthcare. Discover how CorrDyn is helping organizations turn data into breakthroughs at CorrDyn.

About

Data in Biotech is a fortnightly podcast exploring how companies leverage data to drive innovation in life sciences.  Every two weeks, Ross Katz, Principal and Data Science Lead at CorrDyn, sits down with an expert from the world of biotechnology to understand how they use data science to solve technical challenges, streamline operations, and further innovation in their business.  You can learn more about CorrDyn - an enterprise data specialist that enables excellent companies to make smarter strategic decisions - at www.corrdyn.com

You Might Also Like