Intelligent Founder AI Podcast

@poonamparihar

One focused investigation per week: a critical AI or tech story breaking now, what actually changed beneath the headlines, and how to respond for real business value. www.intelligentfounder.ai

  1. Aug 31

    Ep.020 - “How to stop an AI trial becoming a dead end!

    I was OOO all last week and just got back, so I did not publish the deep dive I had been preparing. I’ll finish it and publish two long-form pieces this week. This episode develops the commercial-layer argument from the previous episode which is - a successful technical test is not automatically a successful business outcome. If you have not heard Episode 19, it is worth listening to alongside this one. A Pilot Is Not Progress Until It has a path to deployment. McKinsey’s 2026 global survey found that nearly nine in ten respondents use AI regularly in at least one business function, and 44% say AI is now scaling across their enterprise. Yet only 37% report any positive impact on EBIT, unchanged from the prior year. Just 6% qualify as AI “high performers”: organisations attributing at least 5% of EBIT to AI and describing its effect as significant. That gap is the real AI story. Companies are not short of demonstrations, experiments, workshops, or prototype ideas. They are short of deployments that change a meaningful business result, and of a disciplined route from a limited test to normal operation. For a founder selling AI, a request for a pilot should therefore be encouraging, but never automatic cause for celebration. A pilot can open a door. It can also consume three months of product work, customer support, data cleaning, meetings, and bespoke requests, only to end with: “Very interesting. We will come back to you.” and therefore, a successful test is not the same thing as a commercial success. Intelligent Founder AI is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. The uncomfortable numbers The evidence is mixed in scale, but consistent in direction: the move from experimentation to financial impact remains difficult. Deloitte reports that worker access to AI rose by 50% in 2025. Two-thirds of organisations report productivity and efficiency gains, 53% report better insights and decisions, and 40% report lower costs. But only 20% say they have achieved revenue growth from AI, even though 74% hope to do so in future. An MIT NANDA analysis, reported by Fortune, reached an even sharper conclusion. It found that only around 5% of enterprise generative-AI pilots in its dataset produced rapid revenue acceleration; most delivered little or no measurable profit-and-loss impact. The research drew on 150 executive interviews, a survey of 350 employees, and 300 public deployments. The exact percentage will vary by sector, use case, and how success is defined. But founders should not take comfort in a pilot simply because the customer agreed to run one. The market is full of pilots. What is scarce is a clear decision to deploy, pay, renew, and expand. Beginning at the end - The most useful question to ask before a pilot starts is not, “What can we test?” It is: “If this works, what happens next?” A vague answer like “Let’s see”, does not mean the customer is unserious. It does mean the work is probably a learning exercise, not yet a buying process. A stronger answer would have a chain behind it. It’ll identify the outcome that matters, the person accountable for it, the proof required, the budget route, and the next decision. For example: “If this reduces manual inspection time by 25%, works with our existing reporting process, and completes the security review, our operations director will decide whether to fund deployment at two sites.” That is not a guaranteed deal. but it’s much better: it is a visible path to one. This is particularly important for deep tech and physical AI, where it must operate reliably in a real setting, fit into human routines, cope with imperfect data and connectivity, and give an organisation confidence that it can support the system safely over time. Five questions before you say yes! A worthwhile pilot does not need a giant programme plan. But it should answer five basic questions in plain language. 1. What problem are we solving? “We want to explore AI” is not a problem. “Our engineers spend six hours each week manually reviewing inspection data” is. A strong pilot begins with a pain that someone experiences today. 2. What will improve, and how will we know? Agree a starting point. It could be time per task, false alarms, missed defects, response time, rework, downtime, cost, or revenue. Model accuracy may matter, but it is rarely the business case on its own. 3. Who owns the result? Every pilot needs a named customer-side owner: someone who has an operational reason to care, can bring the right people together, and will still be involved when it is time to decide what happens next. 4. What must be true for deployment? Surface the practical work early: data access, integration, cyber security, legal review, procurement, user training, support, and the workflow for handling uncertainty or mistakes. You may not solve all of it in the pilot, but you should not discover it only after the pilot succeeds. 5. What is the next commercial decision? Name the likely next step - a paid deployment, a defined expansion, an integration phase, or a formal investment case. Name who makes that decision and when. These questions protect both sides. The customer avoids an attractive proof of concept that cannot be used, and the founder avoids turning a product company into an unpaid custom-development team. Building for changed work, and not AI theatre The strongest founders do not merely add AI to an unchanged process. McKinsey found that nearly three-quarters of AI high performers report fundamentally redesigning workflows around AI, compared with only one-quarter of other respondents. That is the point a pilot must test. Not just: “Does the model produce a useful answer?” But: Who receives that answer? What do they do differently? When do they override it? How does the decision get recorded? What measurable operational result follows? Deloitte reaches a similar conclusion: only 34% of organisations are deeply transforming their business with AI, while 37% use it at a surface level with little or no change to existing processes. The difference is not mostly about having access to a better model. It is about ownership, workflow design, data readiness, and the willingness to make a decision once evidence arrives. The founder’s test Before beginning your next pilot, write one sentence that starts with: “At the end of this pilot, the customer will decide whether to…” If you cannot finish that sentence clearly, pause. Ask more questions. Reduce the scope. Find the real operational owner. Agree what success means. Or Describe the work honestly as a learning engagement, price and resource it accordingly, and protect your roadmap. A pilot is not a trophy. It is a bridge. Its value lies not in proving that your technology can work, but in creating enough operational and commercial confidence for a customer to make the next decision. Intelligent Founder AI is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Pilot-to-deployment checklist Use this before you say yes to the next “let’s try a pilot” conversation. It will not remove uncertainty, and it should not really, but “It will help you expose the assumptions early, get the right people in the room, and make sure the pilot begins with a clear decision in mind. * A specific operational problem, expressed in current cost, delay, risk, or workload * A baseline and agreed success measure * A named operational owner and executive sponsor * Defined data, site, integration, security, and user-access requirements * A bounded scope, timeline, responsibilities, and change-control process * A decision meeting booked before the pilot begins * A pre-agreed next step: paid deployment, expansion, integration, or closure * A deployment budget owner, procurement route, and indicative commercial model The point is not to avoid pilots. Good pilots are how customers build confidence in a new capability, and how founders learn what it really takes to operate in the field. But a pilot should do more than produce a promising result or a good case study. It should help the customer make a decision: deploy, expand, integrate, or stop. So, before you commit the team, ask one straightforward question: “If we hit the agreed success criteria, who decides what happens next, where does the budget come from, and when will that decision be made?” If nobody can answer yet, that does not automatically mean walk away. It probably means you are not discussing a deployment pilot. You are discussing discovery. Narrow the work, charge for it, protect your roadmap, and use the engagement to create a clearer path to a real commercial decision. This article accompanies Episode 20 of the Intelligent Founder AI Podcast: “Make the Pilot Count: How to stop an AI trial becoming a dead end.” The episode explores the practical founder playbook for turning a test into a route to deployment. Thanks for reading Intelligent Founder AI! This post is public so feel free to share it. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.intelligentfounder.ai/subscribe

  2. Aug 24

    Ep.019 - “Who Really Buys AI?

    Before we jump into a new series, I want to take a second and reflect. We wrapped a good run of episodes that went deep into infra decisions - compute, models, sovereignty, security, the edge, the whole stack. in total we have 75 total articles on the blog including the deep dives and good numbers of timely takes on news on AI Unfiltered as well, covering length and breath of technology side. What we haven’t yet spoken much about is the far messier, far more decisive question sitting right behind all of it. and that is, how does an enterprise actually select an AI or deep-tech product, sign off on it, deploy it without it dying in a pilot purgatory, renew it, and eventually expand it across the organization? That's the commercial operating layer. this is the layer that determines whether any of the technical decisions All the architecture and model choices in the world don’t matter if nobody inside the enterprise actually says yes to buying the thing. So before we start the next series, I wanted to slow down for an episode or two and talk about that how AI actually gets bought, deployed, and renewed inside a real company. Which brings us to today’s question: who really buys AI? The AI Deal Usually Stalls After the Demo. Intelligent Founder AI is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Enterprise sales are less about persuasion than understanding how a decision gets made. An AI demo can generate real excitement. A team sees a capability working, begins to imagine its potential, and asks to explore a pilot. That is a useful moment. but it’s still not a deal, YET. The gap between interest and purchase is where many AI products stall. The reason is rarely that nobody liked the technology. More often, the organisation has not yet answered the practical questions behind any meaningful purchase: who owns the outcome, where the money comes from, how the tool fits existing work, what data it needs, and who accepts responsibility if it fails. and for founders, this changes the job, because you are not simply trying to convince someone that the product is impressive. You are helping a company make a decision it can stand behind. The hidden work behind a “yes” In an enterprise, enthusiasm can start a conversation, but it cannot complete one. A manager who wants the product may not own a budget. A budget holder may want evidence of return. Technology teams may need to assess integration. Security, data, legal, or procurement teams may need to establish whether the supplier and solution are acceptable. This does not necessarily mean enterprise buyers are slow for no reason. They are managing consequences that do not appear in a demo: operational disruption, data exposure, unplanned implementation work, contractual liability, and the cost of supporting another system, and the founder who understands this early can design a much better sales process. Instead of asking only, “Did they like it?”, we need ask a more useful question: “What needs to be true for this to become a normal part of how you work?” The answer will reveal far more than a positive reaction ever can. Look for a decision, not applause A serious opportunity has a visible path from problem to action. There is a specific operational issue, rather than a broad wish to “do something with AI.” Someone is affected by that issue often enough to care. A business leader can connect improvement to money, time, risk, quality, or growth. The people who will use the product have a reason to adopt it. And there is a realistic route through the company’s technical and commercial checks. When those pieces are missing, a founder can still run a useful experiment. But it should be treated as learning, not forecast as revenue. This distinction becomes especially important when customers request pilots. A pilot can create evidence and confidence. It can also become a holding pattern: lots of effort, plenty of positive comments, and no clear owner for the next decision. Before the work begins, it is reasonable to ask what a successful result would unlock. The answer does not need to be a contract. It should show that the customer has considered the next move. The founder’s job is translation Different people inside a company judge the same product through different lenses. The business sponsor wants a worthwhile outcome. The person doing the work wants the tool to make life easier, not add another dashboard. The technical team wants clarity about how it connects and operates. Risk teams need boundaries, documentation, and accountability. Procurement needs a workable way to buy. The founder’s task is to translate one product into terms that matter to each of these audiences. That is not a matter of changing the truth or inventing different stories. It is about making the same value understandable in the context of each person’s responsibilities. The strongest early customers are often not the largest names. They are the organisations where the problem is sharp, the internal owner is committed, and enough people can move in the same direction. here’s a quick test for your next opportunity! Take one promising customer conversation and answer these questions: 1. What costly or risky thing is happening today? 2. Who is accountable for improving it? 3. What proof would make a wider rollout sensible? 4. What needs to fit technically and operationally? 5. Who can approve the spend? 6. What happens immediately after a successful pilot? If you cannot yet answer most of them, you have more discovery to do. That is not failure. It is the work that turns a promising conversation into a real opportunity. A strong demo opens the door. Understanding how the customer makes decisions is what keeps it open. Episode 19 of the Intelligent Founder AI Podcast explores this in more depth: “Who Really Buys AI? Why a great demo is not enough to win an enterprise customer.” Listen for the practical framework for mapping the people, risks, and decisions behind an enterprise AI purchase. Thanks for reading Intelligent Founder AI! This post is public so feel free to share it. latest on AIUnfiltered and Quantopinion. Intelligent Founder AI is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.intelligentfounder.ai/subscribe

  3. Aug 16

    Ep.018 - Nvidia and the Compute Wars: What It Means If You Are Building Right Now

    Over the past eight episodes, we set out to make one idea practical: building with AI is no longer just about choosing a model or writing a prompt. It is about making a sequence of technical, commercial and operational decisions that determine whether an AI product becomes a durable business. We explored the questions founders now have to answer in real time: where AI genuinely creates an advantage, how to choose between building and buying, how to think about agents and automation, what it takes to move from prototype to production, and why infrastructure choices quietly become product strategy. We did not approach these topics as a catalogue of tools. We approached them as decisions and as trade-offs. This is the final episode of this first run. It closes where many AI businesses ultimately arrive: at compute. Intelligent Founder AI is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Compute can sound like a technical footnote. In practice, it shapes model choice, unit economics, latency, deployment geography, data governance, reliability, fundraising narratives and the speed at which a team can experiment. If the previous episodes were about deciding what to build with AI and how to turn it into a useful product, this episode is about the physical and economic system underneath those choices. What changed? When I first prepared this episode, the story looked relatively straightforward. Nvidia’s data-centre revenue had reached $51.2 billion in a single quarter, Blackwell demand appeared intense, and GPU rental pricing seemed likely to become steadily cheaper as more capacity arrived. The first observation has become even more significant. Nvidia subsequently reported $75.2 billion in data-centre revenue for its fiscal first quarter of 2027, up 21% quarter on quarter and 92% year on year. This is not simply a company earnings statistic. It shows the scale at which accelerated compute has become foundational infrastructure for the AI economy. The second observation also holds: demand for advanced AI systems remains very strong. But the useful founder lesson is not to treat any single backlog figure or “sold out” headline as a permanent fact. Capacity differs by configuration, region, customer relationship, commitment length and workload profile. Supply chains change. customer priorities change; a chip that is scarce for frontier training may be entirely unnecessary for a production inference workload. The third observation needs more nuance. Compute has not followed a smooth, one-way path towards lower prices. Older-generation hardware, marketplace supply, reserved commitments and specialist clouds can create excellent economics. At the same time, premium capacity can tighten quickly when large training runs, agentic products or new model releases absorb supply. The relevant question is therefore not, “Will GPU prices fall?” It is, “What is the least expensive reliable way to meet our specific quality, latency, privacy and volume requirements?” The real Nvidia moat. Nvidia’s position is about more than high-performance GPUs. It is a full operating environment: CUDA, libraries, frameworks, optimized kernels, deployment tooling, networking, systems expertise and a deep global base of engineers who already know how to use it. That matters because switching compute platforms is not the same as changing a cloud region. A model may technically run elsewhere while requiring new optimisation work, different kernels, revised serving infrastructure, a changed observability stack and a fresh performance-validation cycle. The strongest lock-in often lives in engineering time and accumulated operational knowledge, not a contractual restriction. But dominance does not mean every founder should default to the newest Nvidia hardware. It means Nvidia is the benchmark against which other choices should be evaluated. Competition is becoming useful AMD remains the most important broad GPU alternative, while custom silicon from the major cloud platforms has become relevant for teams operating within their ecosystems. Google TPUs can be compelling for selected training and inference workloads. AWS Trainium and Inferentia can offer meaningful savings for compatible models, although savings are not automatic: they depend on model support, migration effort, deployment scale and actual utilisation. AWS has published customer examples of around 50% cost reduction after migration to Inferentia, which should be treated as evidence of potential, not a universal pricing promise The practical consequence is healthy competition, but not effortless portability. Founders should not build around a theoretical ability to run identically everywhere. They should build practical optionality: the ability to test a second credible path before dependency becomes expensive. That means separating application logic from infrastructure-specific code where possible, maintaining reproducible evaluation suites, recording real cost and latency data, and avoiding optimizations that only make sense for a single supplier until there is a demonstrated payoff. Local AI is real! I initially referred to “RTX Spark,” which was incorrect. Nvidia DGX Spark, a compact desktop system built around the Grace Blackwell platform, with 128 GB of unified memory is the new product Nvidia is positioning for local AI development and says it can support models up to 200B parameters, with up to one petaFLOP of FP4 AI performance. those specifications do not translate directly into every model’s real-world inference speed but the broader point stands. Local AI is becoming a genuine part of the development and deployment landscape. For many founders, its value will not be replacing cloud infrastructure. Its value will be faster private experimentation, lower friction during development, work with sensitive data, edge deployment prototypes and the ability to validate a workflow before committing to recurring cloud spend. For regulated, industrial or data-sensitive applications, this is especially important. The architecture may not be “cloud versus edge.” It may be local development, cloud-scale training, and targeted edge or on-premises inference in production. The build-rent-buy staircase A durable compute strategy is not a single decision. It is a staircase. At the bottom is rent: use APIs, managed model services, or on-demand compute while testing whether users care. This keeps capital commitment low and lets the team change direction quickly. The next step is optimize: select smaller or more efficient models, use batching, caching, routing and quantization, and measure whether latency, quality and cost actually meet the product requirement. Most teams can gain more here than by chasing the next hardware release. After that comes commit: reserved instances, capacity agreements, dedicated infrastructure or a managed deployment partner. Move here only when demand is sufficiently predictable that the economic benefit outweighs the loss of flexibility. Finally, for a limited set of high-volume, sensitive or latency-critical workloads, there is own or operate: dedicated hardware, on-premises systems or edge deployment. This can be compelling, but only once the operational burden and lifecycle costs are understood. The mistake is to climb this staircase too early. The opposite mistake is to remain at the expensive, unmeasured rental stage after usage has become stable and large enough to justify change. What the series concludes. Across these eight episodes, the recurring message has been simple: the technology moves quickly, but the disciplines of building an enduring business change much more slowly. * Start from a sharp customer problem rather than a model capability. * Test the workflow before scaling the architecture. * Treat data quality, evaluation and reliability as product work. * Be realistic about integration and human adoption. * Measure unit economics early. And preserve enough technical flexibility that a supplier, model or platform change does not force a rewrite of the business. The compute wars do not remove those responsibilities. They make them more important. For founders, the opportunity is not to predict which accelerator wins. It is to build a product that can benefit from competition without being trapped by it. Use the hardware that delivers the required outcome today. Keep one credible alternative visible. Let evidence, and not hype, benchmark headlines or fear of missing out, determine when to change your stack. AI may or may not company-building easy, but with a clearer view of the choices will make it more possible. Listen to the full episode here, in Substack app, or Apple, Spotify / youtube. Thanks for reading Intelligent Founder AI! This post is public so feel free to share it. New On the Quantopinion and AIU This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.intelligentfounder.ai/subscribe

  4. Aug 7

    Ep.017 - The Edge AI Question

    In our last deep dive, “AI Workloads at the Telco Edge,” we looked at AI inference as a workload-placement problem. Rather than treating “the edge” as a single place, it broke deployment into infrastructure tiers, from regional data centres and metro nodes to enterprise sites and far-edge environments, each with different constraints around latency, bandwidth, power, cooling, and cost. This podcast takes the next step. Instead of asking whether AI should move to the edge, it asks a more useful question: which workloads genuinely need to run close to where data is created, and which are better handled in larger, more economical data centres? This episode is part of my nine-part series, Build vs Buy vs Rent: The AI Infrastructure Decision Tree for Startups. I recorded it a few weeks ago, before publishing my recent deep dive into AI inference infrastructure, but it connects directly to that work by examining the practical reasons organisations choose to run inference locally in the first place. The compliance and network reality and - why data residency, security, connectivity, and operational control can drive edge adoption. Intelligent Founder AI is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Why organisations choose to run AI closer to the data Two forces push AI workloads toward the edge, and they’re often conflated. The first is speed: when a system is controlling equipment, supporting a safety-critical decision, monitoring a live process, or reacting to a real-time physical event, waiting on a round trip to a distant cloud simply isn’t practical. The second, less discussed, force is constraint. Data often has to stay local because of privacy, data residency, security, commercial sensitivity, or unreliable connectivity. A factory, hospital, transport operator, utility provider, or critical-infrastructure site may not be comfortable — or even legally permitted, to send every video stream, sensor reading, or operational record to a central cloud. Together, these forces reframe the edge-AI question. It stops being about chasing the lowest possible latency and becomes about balancing three competing needs: keeping sensitive or high-volume data close to where it’s created, running models reliably where they’re needed, and managing the real cost, power, and operational burden of local compute hardware. Get this balance right, and organisations gain faster decisions, lower bandwidth use, more resilience during connectivity problems, and stronger control over sensitive information, without the assumption that everything must move to the cloud by default. None of this is free, though. AI hardware needs power, cooling, physical space, maintenance, and careful monitoring, and the more capable the model, the greater those demands become. That’s why the real decision isn’t “cloud versus edge” - it’s deciding what must stay local, what can run nearby, and what can still be handled centrally. The business case for moving inference to the edge For founders building AI for physical environments such as manufacturing, transport, energy, agriculture, and retail, a cloud-first default can become the wrong architecture when workloads are persistent, latency-sensitive, connectivity-constrained, or subject to local data requirements. for example lets consider an illustrative deployment with 1,000 inferences per hour across 500 locations. A cloud-based approach can become expensive when large volumes of sensor, image, video, or event data must be processed continuously across hundreds of locations. By contrast, an edge deployment shifts more of that cost into upfront hardware, local power, device management, maintenance, and lifecycle support. The break-even point varies significantly by model size, inference frequency, batching, bandwidth, hardware choice, energy pricing, support requirements, and expected device lifetime. But for sustained, high-volume workloads, the total cost of ownership can favor edge or hybrid architectures over routing every inference through a cloud service. The exact numbers depend on the model, inference volume, cloud pricing, hardware lifetime, local power costs, maintenance, and utilisation, but the underlying shift is what matters: recurring API spend becomes owned infrastructure and operating cost. For high-volume, latency-sensitive, or data-sovereign workloads, edge and hybrid architectures can materially reduce five-year costs compared with routing every inference to the cloud. A useful test is to ask whether at least two of these are true: the application needs a consistently low-latency response; each location generates sustained volumes of data or inference requests; data residency, privacy, security, or commercial requirements make cloud routing difficult; or the application must keep running when connectivity is unreliable. New local hardware is also expanding what’s possible outside a central data centre, systems such as NVIDIA’s Jetson range and desk-side AI development systems point toward increasingly capable models running in compact local environments. . But hardware capacity alone doesn’t make an edge deployment viable; power, cooling, model optimisation, network resilience, monitoring, and maintenance still matter just as much. Listen to the full episode here, on Substack app, or Apple, Spotify / youtube. In the coming weeks, I’ll publish few deep dives on the decisions, constraints, and trade-offs that shape practical edge-AI deployments, so this episode and previous AI workload deep dive become part of a wider series on AI inference infrastructure covering it in 360 degrees- from where workloads should run, to the practical realities of deploying and scaling them outside centralized cloud environments, so subscribe to the newsletter. Thanks for reading Intelligent Founder AI! This post is public so feel free to share it. Latest on AIU and Quantum Opinion This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.intelligentfounder.ai/subscribe

  5. Jul 31

    Ep.016 - Sovereign AI and Compliance: Infrastructure as a Regulatory Decision!

    Most ‘sovereign AI’ debates talk about models and tokens. In practice, regulators and procurement teams care first about where your data flows, where your compute runs, and who can turn it off. For any founder building AI in healthcare, financial services, critical infrastructure, or transport etc, the infrastructure decision is a compliance decision. Get it wrong and the economics become irrelevant because you cannot operate. The UK AI Safety Institute, now a statutory regulator, has mandated that any AI system used in critical infrastructure must be capable of functioning for 72 hours without an internet connection. That requirement alone rules out a purely cloud-hosted stack for safety-critical use cases. The practical response is a hybrid architecture with local inference capability for critical functions and cloud AI for non-sensitive workloads. Imagine you’re shipping a rail monitoring system, and the cheapest v1 is a single cloud‑hosted model behind a US‑based API. It passes a pilot, then fails when the operator’s safety case demands 72‑hour offline capability, local failover, and documented data residency. Suddenly you’re rebuilding the stack under regulatory time pressure instead of product strategy. Under GDPR Article 35, any AI system involving high-risk data processing such as large-scale profiling, special category data requires a Data Protection Impact Assessment before deployment. Intelligent Founder AI is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. If your AI model is called through a US-based API, you are transferring data to a third country, which requires an adequacy decision, standard contractual clauses, or another lawful transfer mechanism. If your infra diagram still routes production traffic through that API, your DPIA isn’t finished, it’s telling you that your infrastructure doesn’t match your legal story yet. Many enterprise procurement contracts now require data not to leave the UK or EU, which effectively mandates UK/EU-based cloud providers or on-premises deployment. The UK is deliberately diverging from the EU AI Act, positioning itself as more permissive for deep-tech development. But if your customers are in the EEA, you need to comply with both frameworks simultaneously. Dual-track architecture designed to meet UK and EU requirements from day one is the pragmatic response for any founder planning cross-border commercial deployment. In practice that means at least one deployment path that never leaves UK/EU infrastructure, and a clearly separated path for more permissive markets, so sales is not blocked by your first cross‑border deal. Building for sovereignty costs more upfront. It costs dramatically less when your first regulated enterprise deal requires it, and that conversation comes sooner than most founders expect. this is 7th episode in the series of Build vs Buy vs Rent: The AI Infrastructure Decision Tree for Startups!. Listen to the full episode here, on Substack app, or Apple, Spotify / youtube. Recent Posts on AIUnfiltered and QuantOpinion.ai Thanks for reading Intelligent Founder AI! This post is public so feel free to share it. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.intelligentfounder.ai/subscribe

  6. Jul 25

    Ep.015 - AI FinOps: Stopping the Invisible Cost Leaks!

    Cloud waste is exploding in 2026, and the big culprit is unmanaged AI workloads. Roughly 80% of AI GPU spend is now always‑on inference in production, quietly burning cash in the background. If you make those systems just 20% more efficient, you’re basically giving yourself a 20% margin boost at scale. This is 6th episode in the series of Build vs Buy vs Rent: The AI Infrastructure Decision Tree for Startups with 2 more to go!. Listen to the full episode here, in Substack app, or Apple, Spotify / youtube. Intelligent Founder AI is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. In practice, the money you spend on AI in the cloud tends to leak out in four specific areas – idle compute, oversized models, wasted tokens, and hidden data transfer fees. Idle compute: GPU instances left running between jobs, notebooks never shut down, development endpoints serving zero traffic. The fix is auto-shutdown after two hours of idle time for development environments, and scheduled shutdown overnight for staging. Model over-provisioning: defaulting to a frontier model for every task. A routing layer sending simple queries to small, cheap models can cut inference costs by 40 to 70 percent with no user-visible quality degradation. Token waste: long system prompts, over-retrieved context, over-long generated outputs. At scale these add tens of thousands of pounds per month. and Egress: data transfer costs between regions or providers, typically 80 to 90 dollars per terabyte, adding 20 to 40 percent to apparent compute costs. FinOps basically means “know where every AI pound is going, then fix the leaks.” First, you track every GPU, every AI API call, and every batch of tokens so you can see simple numbers like cost per user and cost per AI request. If it costs you ten pence to run an AI feature and you only charge five pence, the math will never work, no matter how impressive the top-line revenue looks Then you run a 90‑day clean‑up: month one is wiring up the tracking, month two is killing the three biggest sources of waste, month three is locking in discounts and automatic safeguards. You usually spend just a few percent of your AI budget on FinOps tools, and they can earn themselves back within the first quarter. Recent Posts on AIUnfiltered and QuantOpinion.ai GeoIntel Wargamer - How We Built using Qwen Cloud GeoIntel is a multi-agent AI “debate room” for geopolitics, where a Historian, Strategist, and Economist argue over real-world conflicts and then synthesize probability-weighted forecasts grounded in live web search. It’s inspired by how real intelligence teams use red-teaming and structured analysis, but packaged as software you can query like any other AI tool. Let us know your feedback here : does this feel useful, and what kinds of questions or scenarios would you want to throw at a system like this? Thanks for reading Intelligent Founder AI! This post is public so feel free to share it. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.intelligentfounder.ai/subscribe

  7. Jul 19

    Ep.014 - Open Source vs Proprietary AI: The Model Decision That Shapes Your Infrastructure

    The performance gap between open-source and proprietary frontier models has collapsed. DeepSeek V3 offers performance comparable to GPT-4o at 27 cents per million input tokens, compared to roughly 2.50 dollars for GPT-5. DeepSeek’s reasoning model costs 55 cents per million input tokens - 96 percent cheaper than equivalent OpenAI reasoning models. Open-source models now cover approximately 80 percent of real-world enterprise use cases at 86 percent lower cost than proprietary alternatives. The practical spending threshold is 15K pounds per month in API costs. Below that, the engineering overhead of self-hosting open-source models is not worth it. Above it, the economics justify a proper evaluation. Above 100 million tokens per month with adequate engineering capacity, open-source self-hosting is almost always significantly cheaper. Proprietary models give you access to frontier capability with zero deployment overhead. Intelligent Founder AI is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. The trade-offs are vendor lock-in, no control over model behaviour or training data, and complete dependency on a vendor who can change pricing or deprecate model versions on their timeline. Open-source models give you full control, fine-tuning capability on proprietary data, and deployment flexibility including air-gapped environments. The trade-off is that you own the operational complexity. Fine-tuning pays back when you exceed roughly ten thousand requests per day - below that, prompt engineering with a frontier model is more economical. Modern prompt optimisation techniques have been shown to outperform reinforcement learning fine-tuning by 6 to 19% points on benchmark tasks while using up to 35 times fewer compute resources. Fine-tuning wins for very high volume token cost reduction and for hard-to-prompt output formats. The model choice and the infrastructure choice are the same decision viewed from two angles. Get clear on your volume, your compliance requirements, and your engineering capacity, and the right choice usually becomes obvious. This is 5th episode in the series of Build vs Buy vs Rent: The AI Infrastructure Decision Tree for Startups Listen to the full episode here, in Substack app, or Apple, Spotify / youtube. Thanks for reading Intelligent Founder AI! This post is public so feel free to share it. Latest On AIUnfiltered - GenAI Was the Warning Shot, Agentic AI Is the Next Test Different angles, same battlefield: trust, safety, and control over information. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.intelligentfounder.ai/subscribe

  8. Jul 13

    Ep.013 - The On-Premises Case: When Buying Hardware Actually Wins

    If you’re pushing serious AI workloads, there’s a point where buying your own GPUs quietly beats “just use the cloud” on pure maths. Once you cross that point, the savings stop being theoretical and start showing up in your P&L. This is 4th episode in the series of Build vs Buy vs Rent: The AI Infrastructure Decision Tree for Startups. TL;DR * Above ~70% GPU utilisation, owning hardware usually beats the cloud on cost. * The real break‑even sits roughly between 55–75% utilisation, depending on power, amortisation, and cloud pricing. * Lenovo’s TCO work: ~8x cheaper than cloud infra and up to ~18x cheaper than frontier APIs per million tokens. * At 10B tokens/month, three‑year savings vs pure APIs can exceed £2M. * 8x H100 server: $250k–$400k upfront plus $3k–$5k/month to run; ~$11k/month effective cost over three years. * Equivalent cloud H100 capacity: roughly $14k–$20k/month. * You must factor in power (6–10 kW per rack unit), UK colocation (£500–£2,000/rack/month), and infra engineers (£80k–£140k/year). * A hybrid model (own the baseline, rent the spikes) can cut AI infra spend by ~40–60% vs pure cloud. Intelligent Founder AI is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. At around 70 percent GPU utilisation, owning on‑prem infrastructure is usually the better deal for AI inference. Below that level, the cloud’s flexibility earns its premium because you’re not paying for expensive hardware that sits idle when traffic drops. In practice, the break‑even lives somewhere between 55 and 75 percent sustained utilisation, depending on things like your electricity rate, how many years you plan to amortise the kit, and which cloud pricing tier you’re comparing against. Lenovo’s 2026 total cost of ownership work puts real numbers to this. They found that self‑hosted GPU infrastructure can be about 8 times cheaper per million tokens than raw cloud infrastructure, and up to 18 times cheaper than using frontier model APIs. Once you’re at around ten billion tokens a month, the three‑year cost gap between owning hardware and living entirely on cloud APIs can easily exceed two million pounds. The sticker price on an 8‑GPU H100 box is not small. You’re looking at roughly 250,000 to 400,000 dollars upfront for the server itself. Then you add 3,000 to 5,000 dollars a month in operating costs for colocation, power, cooling, and maintenance. Spread the hardware over three years and your effective monthly cost lands at around 11,000 dollars. Buying similar H100 capacity on‑demand in the cloud typically ends up between 14,000 and 20,000 dollars a month, so the saving is real and it compounds over time. Where founders often get caught out is in the hidden line items. Modern H100 servers can pull 6 to 10 kilowatts per rack unit, which means you can’t just drop them into a normal office and hope for the best. You need proper co-location, and in the UK that runs about 500 to 2,000 pounds per rack per month. You also need engineers who can run GPU infrastructure safely and reliably, and UK market rates put that at roughly 80,000 to 140,000 pounds per person per year. On top of that, you’re carrying hardware risk: once you buy, your performance ceiling is locked in for the amortisation period while cloud alternatives quietly keep improving in the background. That’s why most serious AI teams don’t go “all cloud” or “all on‑prem” for long. The model they converge on is hybrid: own enough hardware to cover your predictable baseline workloads, and use the cloud when you need to absorb burst traffic. if planned well, that blended architecture cuts your overall AI infrastructure bill by about 40 to 60 percent compared to living entirely in the cloud. It’s where most AI companies end up once their volume of inference forces them to care about infrastructure as more than a line item. Listen to the full episode here, in Substack app, or Apple, Spotify / youtube. Thanks for reading Intelligent Founder AI! This post is public so feel free to share it. This week Intelligent Founder AI and AI Unfiltered broke into Substack’s “Rising in Technology” leaderboard at #88—together, after already hitting #98 in Business earlier this year. Thank you for all the support so far. AI Unfiltered is where we track the latest AI news, investigations, and “what just happened?” moments and, more importantly, what they actually mean for operators, founders, and buyers. Think model‑on‑model training fights, national‑security angles, regulatory shifts, and how all of that moves the ground under your product roadmap. Intelligent Founder AI goes deeper: long‑form dives into technical and business architecture, operator‑grade GTM, and governance playbooks for getting AI into real enterprises and critical infrastructure without getting lost in the hype. It’s where we turn those headlines into concrete frameworks you can actually run inside your own stack. Together, AIU and IF.ai are designed as a pair: one keeps you on top of what’s changing week to week, the other helps you re‑wire your architecture, strategy, and sales motion around it. If that sounds useful, check out both, hit subscribe, and tell us what you want us to dissect next. Intelligent Founder AI is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.intelligentfounder.ai/subscribe

About

One focused investigation per week: a critical AI or tech story breaking now, what actually changed beneath the headlines, and how to respond for real business value. www.intelligentfounder.ai