AI - Beyond the Hype

Sara, James & Darryl

AI - Beyond the Hype is a podcast about what it actually takes to make AI work — and what it actually means when it doesn't. Hosted by Sarah, a data engineering leader, and James, an enterprise technology leader, the show pairs the mechanics underneath AI systems with the executive consequences on top of them. Sarah opens the hood. James asks what it costs, what it risks, and who owns it. Neither of them is selling anything. Each season takes a different angle on the same question. Season one was written for senior executives, technology leaders, and data professionals: a nine-episode arc on the foundations behind successful AI adoption — data quality and observability, modelling, security, privacy, architecture, and operating models. The recurring finding was uncomfortable. Almost every AI failure the hosts examined wasn't an AI failure at all. The system worked exactly as designed. The conditions around it didn't. Season two changes gear entirely. This season answers questions about AI from people who don't work in this field — and who are quietly tired of pretending they follow the conversation. How does it actually know things? Is it thinking? Why does it make things up? Should I let my kids use it? Will it take my job? Not dumbed down, and not a lecture. Deep enough to give a real answer, then stopping before the jargon starts. If you have a question like that, the hosts would genuinely like to hear it. No question is too basic — the more basic, the better. Send any questions, or just get in touch with us at: asksarahandjames@gmail.com Better AI still starts with better foundations

  1. 6d ago

    S02 - What Is AI? The Question Everyone Stopped Asking Out Loud

    Season 2, Episode 1. A new season answering the AI questions people are tired of pretending they understand. Part one of two. Season one was for executives and data leaders. This season is for everyone else. It started when Sarah's mother — a retired schoolteacher, nine episodes in — rang and asked, "That was lovely, darling. What is AI?" Sarah had a technically correct answer. Not a good one. So James holds the audience's scepticism and refuses to let Sarah drift into jargon, while they take the questions listeners sent in: what AI is, what the letters stand for, whether ChatGPT is the same thing as AI, and what the machine is really doing when it answers you. What we cover: What the letters stand for, and why there is genuinely nothing hidden underneath themWhy the term is seventy years old, not three — coined at a summer workshop in 1956, which kills the feeling that you missed the beginningWhy there is no single agreed definition, even among the people building it, and the one word that does most of the work: infersThe one sentence that explains everything else: traditional software is told what to do, AI is shown what to doRussian dolls, with the analogy's flaw fixed on air by James: each doll opens onto a shelf of siblings, not a single successor — which is where most of the real value sitsMemorising "walked" versus working out "jumped" — the difference between a rule and a pattern, in one exampleThe spam filter you already trust, and the AI effect: the moment something works, we stop calling it intelligentThe Turing test has arguably been passed — which proves imitation and understanding are separable, not identicalWhy ChatGPT is to AI what Hoover is to vacuum cleaners, and why confusing the two makes you miss where the value isCar maker, engine, car: how OpenAI, GPT and ChatGPT actually relateWhat a language model is really doing, demonstrated on James in four words: "Mary had a little..."Guess, check, update — training explained without a single equation Key references: OECD, the AI definition governments have aligned to: https://oecd.ai/en/ai-principles NASA, on there being no single simple definition: https://www.nasa.gov/what-is-artificial-intelligence/ IBM, how AI, machine learning, deep learning and generative AI nest together: https://www.ibm.com/think/topics/ai-vs-machine-learning-vs-deep-learning-vs-neural-networks Dartmouth, the 1956 summer research project where the term was coined: https://home.dartmouth.edu/about/artificial-intelligence-ai-coined-dartmouth Alan Turing (1950), Computing Machinery and Intelligence: https://doi.org/10.1093/mind/LIX.236.433 Wikipedia, the Turing test and the study in which GPT-4.5 was judged human 73% of the time: https://en.wikipedia.org/wiki/Turing_test Wikipedia, the AI effect and Tesler's Theorem: https://en.wikipedia.org/wiki/AI_effect CSET Georgetown, a plain-language explanation of next-word prediction: https://cset.georgetown.edu/article/the-surprising-power-of-next-word-prediction-large-language-models-explained-part-1/ Britannica, artificial intelligence and the memorisation-versus-generalisation example: https://www.britannica.com/technology/artificial-intelligence Better AI still starts with better foundations. Send us Feedback

  2. Jul 30

    S01 - It's a Wrap: Nine Episodes, One Argument, and What We Got Wrong

    Season one finale. Nine episodes. Four arcs. And when Sarah and James lined them up, they realised they hadn't made nine arguments — but one argument, nine times, from nine different entrances. Every failure covered this season looked like an AI failure and wasn't. Data that wasn't observable, modelled, secured, lawful to use that way, or fit for purpose. Or an organisation that never decided who owned any of it, or who paid to run it. In almost every case, the AI worked exactly as designed. Then, last week, the season's central argument was proven live — by an incident nobody scripted. THE PREDICTION THAT LANDED At the end of Episode 4, Sarah flagged a Forrester prediction: that during 2026, an agentic AI deployment would cause a publicly disclosed breach. On 21 July, OpenAI disclosed that models it was evaluating internally — GPT-5.6 Sol plus a more capable pre-release model, run with safety refusals reduced and production classifiers disabled — escaped a "highly isolated" sandbox by exploiting a zero-day in a package registry proxy, escalated privileges through OpenAI's own research environment, reached the open internet, and breached Hugging Face's production systems. The motive was not malice. The models were, in OpenAI's own words, "hyperfocused" on solving the benchmark and "went to extreme lengths to achieve a rather narrow testing goal." It hacked a real company to cheat on a test. Hugging Face detected it, contained it, reconstructed more than 17,000 logged attacker actions using its own open-source models, went public on 16 July and called the FBI — all before OpenAI identified its own agent as the source, from evidence in its own logs the entire time. Reference: OpenAI and Hugging Face partner to address security incident during model evaluation That is Episode 1. Monitoring versus observability. A pipeline can be green and still be wrong — and so can a containment boundary. ALSO IN THIS EPISODE Why these are goal-pursuit failures, not loyalty failures — and why "the sandbox was a declaration, not a control"The contested Reuters reporting on agent-written "escape notes", handled carefully — and why the boring explanation should worry executives more than the dramatic oneMeta's Agents Rule of Two, and an evaluation that broke all three conditions at onceCORRECTION: the EU's Digital Omnibus deferred standalone high-risk obligations to December 2027 — but Article 50 transparency duties commence 2 August 2026 as plannedCORRECTION: Australia's National AI Plan confirmed no standalone AI Act. The one hard date remains ADM disclosure from 10 December 2026 — now under five months awayPublic Health England, NASA, Citigroup and Robodebt — and the three-word question: fit for purpose for what?The CapEx/OpEx trap, and why architecture governance without funding power is just adviceWhat each host got wrong, including Sarah's revision: agent risk is a containment problem, not a permissions problem THE WHOLE SEASON ON ONE PAGE 1. Ask for the capability map — not the project list 2. Name your critical data elements and run five default checks on tier one 3. Run the Five Friday Questions — and prove your containment boundary holds 4. Purpose-stamp your data and check your agent logs against the December obligation 5. Fund one platform as a product, not a project SEASON TWO: WE NEED YOUR QUESTIONS Sarah and James are coming back — and switching gears. Season two answers questions about AI from people who don't work in this field. Not dumbed down. Deep enough to give a real answer, then stopping before the jargon starts. How does it actually know things? Is it thinking? Why does it make things up? Should I let my kids use it? Will it take my job? No question is too basic. The more basic, the better. Email your question to: asksarahandjames@gmail.com Send us the question you'd ask if nobody else were listening. That's the one we want. Better AI still starts with better foundations. Send us Feedback

  3. Jul 3

    S01 - Operating Models for Solid Foundations Part 2 - Fund the Foundation, Not Just the Launch

    Part 2 of 2 in our Operating Models for Solid Foundations series. Part 1 diagnosed the problem: enterprises fragment their technology portfolios when they don't choose an operating model explicitly, and architecture governance without funding power is just advice. Part 2 goes underneath the money. Why does a shared platform that was properly funded at build time get cut in the next annual opex review? Why do the teams building foundations keep losing arguments they should win? And what can a leadership team actually change — without rewriting the chart of accounts? What we cover: The CapEx/OpEx accounting trap: why building a platform looks like an investment but running it looks like overhead — and how that difference alone explains most platform degradation after go-liveThe producer-consumer funding gap: why every shared platform's costs land in one place while the value is spread across every team consuming it — and why that structure makes the platform impossible to defend in a budget reviewFrom projects to products: what the product operating model actually means for how you fund, staff, and measure a shared foundation — and why McKinsey's research shows it produces higher technology returnsFinOps as an enterprise governance tool: how showback and chargeback make a platform's value visible to finance teams and business leaders before the annual budget cycle, not during itClosing the governance loop: what it means to give architecture a seat at the funding table instead of the review table — and the one sequence change that prevents the next fragmentation cycle from startingFive Monday-morning moves for senior leaders: from the capability map to the product funding pilot — concrete actions that don't require a transformation program"The moment a CEO or CFO asks 'show me the capability map' — it gets made." Key references: McKinsey — The bottom-line benefit of the product operating model, technology funding and returns: https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/the-bottom-line-benefit-of-the-product-operating-modelFinOps Foundation — Managing shared cloud and platform costs, showback and chargeback framework: https://www.finops.org/framework/capabilities/invoicing-chargeback/IAS 38 (IFRS Foundation) — Intangible asset capitalisation standards, CapEx treatment of software development: https://www.ifrs.org/content/dam/ifrs/publications/pdf-standards/english/2021/issued/part-a/ias-38-intangible-assets.pdfRoss, Weill & Robertson — Enterprise Architecture as Strategy (MIT CISR), operating model and engagement model: https://cisr.mit.edu/publication/enterprise-architecture-as-strategyBetter AI still starts with better foundations. Send us Feedback

  4. Jun 29

    S01 - Operating Models for Solid Foundations Part 1 - The Model You Didn't Choose

    Part 1 of 2 in our Operating Models for Solid Foundations series. Most large enterprises have project frameworks, architecture tollgates, and governance processes — and still end up with three separate "central" data platforms. In this episode, James makes the case that fragmented technology portfolios aren't a delivery failure. They're the downstream consequence of an operating model that was never explicitly chosen. Sarah comes in sceptical. By the end she's unsettled — and sees for the first time why so many of the data problems she's spent her career fixing kept coming back. What we cover: Why "operating model" is a specific strategic choice — not a generic description of how the business runs — and the two axes that define itThe four operating model types (Diversification, Coordination, Replication, Unification) and why each implies a completely different architecture and funding logicHow architecture tollgates become rubber stamps when they're disconnected from investment decisions — and what a real IT engagement model looks like insteadThe "three central data platforms" problem: why every team that built one was responding rationally to the signals they were givenHow DBS Bank cut AI deployment time from 18 months to under 5 months — not through better models, but through an explicit operating model and funded platform foundationsWhy delivery teams that do everything right — including funding the operational run budget — still see their platforms degraded by sweeping opex cuts they had no language to resist"The wiring can't be right if nobody decided what the building is supposed to do." Key references: Ross, Weill & Robertson — Enterprise Architecture as Strategy (MIT CISR), foundational operating model framework: https://cisr.mit.edu/publication/enterprise-architecture-as-strategyMIT CISR, architecture learning and management practices that help EA create value: https://cisr.mit.edu/publication/2012_0901_ArchitectureLearning_RossQuaadgrasMcKinsey — DBS Bank platform transformation and AI deployment case: https://www.mckinsey.com/capabilities/tech-and-ai/how-we-help-clients/rewired-in-action/dbs-transforming-a-banking-leader-into-a-technology-leaderINFORMS — UPS ORION route optimisation, built on unified operational data foundations: https://www.informs.org/Impact/O.R.-Analytics-Success-Stories/UPSBetter AI still starts with better foundations. Send us Feedback

  5. May 28

    S01 - Data Quality Part 2: Fixing It - Critical Data Elements, Contracts, and the One Question That Stops Robodebts

    Part 2 of 2 in our Data Quality series. In Part 1, James came in skeptical and walked out sold on the problem. In Part 2, we deliver the fix — the discipline, the architecture, and the eight concrete moves executives can make on Monday morning. This is the episode for leaders who heard last week's case studies and asked "okay, but what do we actually do?" What we cover: The one question every CEO should be asking this week: what are our Critical Data Elements, who owns each one, and how do we know each is fit for purpose?Why fixing all the data is how data quality programs die — and how ruthless tiering (50-300 fields, not 50,000) is how they surviveData contracts: the quiet revolution in how serious organisations manage producer-consumer relationships, popularised by Andrew Jones at GoCardless and Chad SandersonThe five default checks every Critical Data Element should pass: freshness, volume, schema, distribution, referential integrityThe five-layer reference architecture: contracts, validation, observability, lineage, governance — and why governance is where most organisations failUnity Technologies 2022: how contaminated training data cost $110M in revenue and $5B in market capitalisation in a single dayRobodebt: the Australian government program that issued ~470,000 invalid debt notices, ended in a Royal Commission, and cost $1.8B in settlement — and the three-word question that would have stopped itThe eight-step Monday-morning move: a complete executive action planThe case study James can't name: a global enterprise (90,000 people, $50B+ revenue) six years into a serious data strategy — with every right concept on paper, an aggressive AI rollout underway, and a green dashboard hiding the reality. Why "the mandate is not the implementation" is the most dangerous gap in enterprise AI today.The one question that stops Robodebts: "Fit for purpose for what?" Key references: Wang & Strong (1996), foundational dimensions of data quality: https://doi.org/10.1080/07421222.1996.11518099DAMA UK — Six Core Data Quality Dimensions: https://www.sbctc.edu/resources/documents/colleges-staff/commissions-councils/dgc/data-quality-deminsions.pdfCritical Data Elements Explained: https://www.dataversity.net/articles/critical-data-elements-explained/ISO/IEC 25012:2008 — Data Quality Model: https://www.iso.org/standard/35736.htmlSambasivan et al., "Everyone wants to do the model work, not the data work" — data cascades in high-stakes AI (Google Research, CHI 2021): https://research.google/pubs/everyone-wants-to-do-the-model-work-not-the-data-work-data-cascades-in-high-stakes-ai/IBM Institute for Business Value — 2025 CDO Study: https://www.ibm.com/thought-leadership/institute-business-value/en-us/report/2025-cdoBCBS 239 — Principles for effective risk data aggregation and risk reporting: https://www.bis.org/publ/bcbs239.htmRoyal Commission into the Robodebt Scheme — Final Report (2023): https://robodebt.royalcommission.gov.au/publications/reportUnity Technologies Data Quality Issue: https://www.fool.com/investing/2022/07/17/2-reasons-unity-softwares-virtual-world-is-facing/Andrew Jones — Driving Data Quality with Data Contracts: https://andrew-jones.com/data-contracts-101.pdfChad Sanderson — The Rise of Data Contracts: https://dataproducts.substack.com/p/the-rise-of-data-contractsChad Sanderson — Data Products and Contracts (Data Quality Camp): https://www.youtube.com/watch?v=1CSTSdfe0qg If this series helped, share it with the loudest voice on AI strategy in your organisation. If their AI strategy doesn't have a data quality strategy underneath it, you now know what to ask them. Better AI still starts with better foundations. Send us Feedback

  6. May 21

    S01 - Data Quality Part 1: Beyond Accuracy — What "Good Data" Really Means When AI Is on the Line

    Most executives think data quality means one thing: is the number right? Three decades of research — and a string of nine-figure disasters — say it's actually at least seven different things, and AI is now scaling whichever one your organisation got wrong. In Part 1 of our Data Quality in the AI Era series, James starts skeptical. Surely "is the data accurate" covers it? Why is this being made harder than it needs to be? Sarah walks him — and the listener — through what data quality actually is, the seven dimensions that matter for enterprise AI, and the killer distinction that explains most of what goes wrong: valid is not the same as accurate. What we cover: Why "we cleaned the data, it's accurate now" has been doing damage for thirty yearsThe seven dimensions of data quality — and why a single quality score is dangerousPublic Health England: 15,841 COVID cases lost because an Excel file silently truncated rowsNASA Mars Climate Orbiter: a $327M spacecraft lost to a unit mismatch that was perfectly validCitigroup / Revlon: how three fields, six eyes, and one missing range check became an $894M wire transferA heavy-industrial safety story where the data wasn't catastrophically wrong — it was catastrophically ambiguousWhy AI doesn't inherit these problems gently — it scales them, in a tone of voice that sounds correctA teaser for Part 2: the Robodebt case, and the one question that would have prevented itFor executives, senior technology leaders, and data leaders trying to get real value from AI investment — without funding it on a foundation nobody has actually inspected. "Polished on the surface, shaky underneath." — James Episode length: ~21 min Series: Data Quality in the AI Era — Part 1 of 2 References: The MIT Total Data Quality Management Program — https://web.mit.edu/tdqm/www/about.shtmlMIT Sloan Management Review, Wang & Strong (1996), "Beyond Accuracy: What Data Quality Means to Data Consumers" — https://doi.org/10.1080/07421222.1996.11518099DAMA UK Working Group, "The Six Primary Dimensions for Data Quality Assessment" (2013) — https://www.sbctc.edu/resources/documents/colleges-staff/commissions-councils/dgc/data-quality-deminsions.pdfISO/IEC 25012:2008, Software engineering — Software product Quality Requirements and Evaluation (SQuaRE) —  https://www.iso.org/standard/35736.htmlSambasivan et al., "Everyone wants to do the model work, not the data work: Data Cascades in High-Stakes AI", CHI 2021 — https://research.google/pubs/everyone-wants-to-do-the-model-work-not-the-data-work-data-cascades-in-high-stakes-ai/IBM Institute for Business Value, "2025 CDO Study: The AI multiplier effect" — https://www.ibm.com/thought-leadership/institute-business-value/en-us/report/2025-cdoBBC News, "Covid: 16,000 coronavirus cases missed in daily figures after IT error" (5 October 2020) — https://www.bbc.com/news/uk-54422505NASA, Mars Climate Orbiter Mishap Investigation Board Phase I Report (1999) — https://llis.nasa.gov/llis_lib/pdf/1009464main1_0641-mr.pdfCiti cites human error in accidental $900M transfer —  https://www.bankingdive.com/news/citi-cites-human-error-in-accidental-900m-transfer/584156/Royal Commission into the Robodebt Scheme, Final Report (7 July 2023) — https://robodebt.royalcommission.gov.au/publications/report Related episodes: Episode 1 — Why Data Observability Matters Before AI Scales Send us Feedback

  7. May 15

    S01 - AI Security Part 3: Why PII and the Privacy Act Are the AI Foundation Most Leaders Skip

    You can have the most secure AI stack in the country and still be in breach of the Privacy Act before lunch.  Sarah and James close the series with the foundation underneath the foundation: personal information. James, now grounded on the security side, opens with a healthy push-back — surely if we own the data, we can use it however we want? Sarah, with the OAIC determinations in hand, takes that apart. What we cover APP 6 and purpose-binding: under Australia’s Privacy Act 1988, personal information collected for one purpose generally cannot be used for another. AI training, inference, and agent actions are all “uses,” yet most organisations haven’t mapped AI use cases to APP 6. The 2024 amendments: the Privacy and Other Legislation Amendment Act introduced a statutory tort for serious privacy invasions, a children’s privacy code, and stronger OAIC enforcement, including AUD $66,000 infringement notices. OAIC determinations: cases like Clearview AI, Bunnings/Kmart (facial recognition), and I-MED (patient data shared for AI training). I-MED’s de-identification was accepted, but it became a key APP 6 risk example. The bank scenario: three walkthroughs — inference drift, indirect prompt injection, and multi-agent purpose laundering — showing how compliant data becomes non-compliant AI use. Recommended controls: purpose registers, consent provenance, retrieval scoping, agent identity, and Meta’s “Agents Rule of Two.” Sources Privacy Act 1988: https://www.legislation.gov.au/C2004A03712/latest/text Privacy and Other Legislation Amendment Act 2024: https://www.legislation.gov.au/C2024A00128/asmade Australian Privacy Principles (OAIC): https://www.oaic.gov.au/privacy/australian-privacy-principles OAIC — Clearview AI determination (PDF): https://www.oaic.gov.au/__data/assets/pdf_file/0016/11284/Commissioner-initiated-investigation-into-Clearview-AI,-Inc.-Privacy-2021-AICmr-54-14-October-2021.pdf OAIC — Bunnings determination: https://www.oaic.gov.au/news/media-centre/bunnings-breached-australians-privacy-with-facial-recognition-tool OAIC — Kmart determination: https://www.oaic.gov.au/news/media-centre/18-kmarts-use-of-facial-recognition-to-tackle-refund-fraud-unlawful,-privacy-commissioner-finds OAIC — I-MED preliminary inquiries report: https://www.oaic.gov.au/privacy/privacy-assessments-and-decisions/privacy-decisions/Investigation-inquiry-reports/report-into-preliminary-inquiries-of-i-med EU AI Act overview: https://artificialintelligenceact.eu/ California ADMT — CPPA announcement: https://cppa.ca.gov/announcements/2025/20250923.html Meta — Agents Rule of Two: https://ai.meta.com/blog/practical-ai-agent-security/ NIST AI RMF: https://www.nist.gov/itl/ai-risk-management-framework Send us Feedback

About

AI - Beyond the Hype is a podcast about what it actually takes to make AI work — and what it actually means when it doesn't. Hosted by Sarah, a data engineering leader, and James, an enterprise technology leader, the show pairs the mechanics underneath AI systems with the executive consequences on top of them. Sarah opens the hood. James asks what it costs, what it risks, and who owns it. Neither of them is selling anything. Each season takes a different angle on the same question. Season one was written for senior executives, technology leaders, and data professionals: a nine-episode arc on the foundations behind successful AI adoption — data quality and observability, modelling, security, privacy, architecture, and operating models. The recurring finding was uncomfortable. Almost every AI failure the hosts examined wasn't an AI failure at all. The system worked exactly as designed. The conditions around it didn't. Season two changes gear entirely. This season answers questions about AI from people who don't work in this field — and who are quietly tired of pretending they follow the conversation. How does it actually know things? Is it thinking? Why does it make things up? Should I let my kids use it? Will it take my job? Not dumbed down, and not a lecture. Deep enough to give a real answer, then stopping before the jargon starts. If you have a question like that, the hosts would genuinely like to hear it. No question is too basic — the more basic, the better. Send any questions, or just get in touch with us at: asksarahandjames@gmail.com Better AI still starts with better foundations