AI to ROI

Ray Rike

AI to ROI is a podcast that shares how enterprises translate AI investments into measurable business value. Hosted by Ray Rike, Founder and CEO of Benchmarkit, the show features senior enterprise leaders and AI software executives who share how AI initiatives move from pilots to production, and how ROI is actually measured and achieved. In addition, each week, we publish a bonus episode with AI to ROI Newsletter co-author, Peter Buchanan to discuss the Big Story of the Week. The AI to ROI podcast is the evolution of the original "Metrics to Measure Up" podcast.

  1. 2d ago

    Can U.S. Frontier AI Labs Survive a Price War with Open-Weight Models?

    When Kimi K3 landed, the headlines said Chinese open-weight models had caught the American labs, and the AI trade sold off from chipmakers to the labs themselves. Nobody stopped to ask whether cheaper and more profitable mean the same thing. In this week's AI to ROI Big Story, Ray Rike and Peter Buchanan run the actual business math on both sides of the fight and find that neither side has the balance sheet to fight a sustained price war. The setup is stark. Anthropic is projecting its first-ever quarterly operating profit of roughly $559 million in Q2, with an annualized revenue run rate near $47 billion, up from $9 billion at the end of last year. That profit disappears the moment it tries to match open-weight pricing. Gross margin is running around 40%, roughly 10 points below the internal forecast, against more than $350 billion in data center commitments coming due over the next three to five years. OpenAI's picture is even thinner: roughly $30 billion in ARR, a projected $14 billion operating loss this year, data center commitments approaching $1 trillion, and an advertising business off to a slow start that needs to reach $100 billion by the end of the decade to close the gap. What the episode covers: Why the Chinese open weight labs are not the subsidized price killers the coverage assumed, with Z.ai's gross margin falling from 41% to roughly 15%, DeepSeek near break-even on about $500 million of revenue and already back in market after a $7 billion round, and Moonshot and MiniMax raising at rising valuations rather than running toward profitThe open weight versus open source distinction that changes the entire economic model, since every new customer requires more chips, power, and data center capacity, and Z.ai's own numbers show roughly 49 to 50% gross margin on customer hosted deployments versus about 19% when they host and serve via APIThe price war math itself: frontier models cost roughly $6 to $8 per million output tokens to serve, Anthropic's $25 per million on Opus produces about a 70% gross margin, and repricing down to the $4 to $6 range where Meta's Muse Spark sits flips that margin from positive 70% to negative 65%Why DeepSeek cut prices on a low-end model and then, two weeks later, told customers to prepare for substantial increases across the line, particularly on API accessGoogle as the structural outlier, with 83% growth in its cloud and AI segment, Gemini embedded across fifteen products with more than a billion users each, and the Apple Siri deal extending reach toward two billion devicesWhere Kimi K3 actually fits, including the caveats nobody is pricing in: weights released only last week, no published large-scale production deployments, two to three times the token consumption on complex tasks, and infrastructure requirements around a 72 GPU rack that costs millions to install and millions a year to runThe geopolitical wildcard, with Washington weighing sanctions or outright bans on Chinese open-weight models and distillation-related IP exposure still unresolved The metric that resets the argument: cost per completed task Price per token is the easiest unit to measure and the wrong one to buy on. Ray and Peter walk through a frontier lab evaluation that assumed a fully loaded remediation cost of $17 per failed attempt, roughly 10 minutes of a human operator's time. Claude Opus completed the task about 90% of the time at roughly $2.56. Meta's Muse Spark, priced at a quarter of Opus on tokens, succeeded 75% of the time and landed above $5.80 per completed task. The list price was 75% lower, and the delivered cost was more than double. A fifteen-point reliability gap did all the work. The formula Ray offers turns a squishy quality debate into something a CFO can actually evaluate: cost per attempt, plus failure probability times fully loaded remediation cost, divided by success rate. What CFOs and GTM leaders should take away: Build cost per completed task into vendor evaluation and make vendors compete on that number rather than on a token price listSegment AI workloads by what a failure actually costs, since a wrong answer in a regulated process or a customer support interaction carries a very different price than the model delta suggests, and standardizing on one model across both workload types to save on price is the common mistakeTreat orchestration as a cost lever, not just the model choice, since Cursor's internal testing found coordination across multiple models delivered comparable code quality at a fraction of the cost of a single large modelRun your own evaluations instead of trusting public leaderboards, since LMSYS Chatbot Arena measures human preference rather than task completion and says nothing about your workloadWatch the switching trap, because a cheap model that attracts heavy traffic today can reprice two or three times higher in a quarter once the vendor needs marginDo not overreact to the open weight scare by standardizing on the cheapest option, since compute, power, and people costs are accelerating, and vendor durability still belongs in the evaluation Ray's read: on real-world performance and total cost of ownership, the closed-weight labs still hold the advantage, their prices keep coming down, and they have every incentive to avoid a price war. Peter's read: open weight economics favor whoever hosts the model more than whoever built it, which makes the hyperscalers the quiet winners regardless of how this resolves. For the full analysis behind this week's big story, subscribe to the AI to ROI newsletter at: ai2roi.substack.com See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

    Can U.S. Frontier AI Labs Survive a Price War with Open-Weight Models?
  2. Aug 12

    Web Presence Intelligence with Stephan Bajaio, Co-Founder & CEO, VibeLogic

    On this week's AI to ROI podcast, Ray Rike is joined by Stephan Bajaio, co-founder and CEO of VibeLogic and a former co-founder of Conductor, the enterprise SEO platform he helped build over 14 years, culminating in a WeWork acquisition, a management buyback, and a valuation approaching half a billion dollars. With 25 years spanning e-commerce, Yahoo, Time Inc., and a stint as a software CMO, Stephan brings an operator's view of what has actually changed in the buyer journey and what has not. The conversation centers on a hard number. When Stephan asked the marketing team at a $3 billion company what share of their web traffic was being driven by LLMs, the estimates ranged from 20 percent to 50 percent. The measured answer was 1 percent, with a high conversion rate on that small base. That gap between perceived and measured contribution is the core problem for any executive being asked to fund an AEO, GEO, or AIO program this year. Stephan introduces Web Presence Intelligence, a supply-and-demand framing designed for executive conversations rather than channel specialists. Demand is what goes into the search bar or the prompt. Supply is everything that comes back, including publishers, Reddit threads, affiliates, partners, and competitors. The strategic question becomes whether you are influencing where your buyer's opinion is formed, before you decide where to place the bets. Ray pushes back directly, arguing that understanding where models source citations should now outrank owned web properties. Stephan holds his position, and the exchange gets to the heart of the allocation decision facing CMOs and CFOs. In this episode: Web Presence Intelligence defined, and why supply and demand is the right executive language for a channel conversation The four P's of owned content: point, paragraph, page, path, and the audit finding that most sites never actually name the problem they solve in the customer's own words Why owned assets matter more when the interpretation layer keeps changing, illustrated by a professional who lost a decade of LinkedIn equity overnight Attribution reframed as a consequence of measurement rather than a measurement itself, and how to work backward from the sale to the second and third best proxies The baseline requirement, or what Stephan calls the before photo, and why AI deployed against a process you do not already understand produces outcomes you cannot judge Control groups in practice, including a market share test on terminology that showed one enterprise software company exactly what its brand guidelines were costing it in visibility Why the gold rush money went to the people selling pans, and where the actual near-term return sits: auditing your own sales calls, renewal conversations, and customer logs From funnel to hourglass, and why channel ownership models are breaking down as generalists get more capable with AI Augment, automate, or build net new: three different projects that require three different budgeting methods and three different attribution models The operator takeaway: control what you can control, establish a baseline before you fund the initiative, and structure AI marketing investments so the return can be defended against assets you own rather than placements you do not own. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

    Web Presence Intelligence with Stephan Bajaio, Co-Founder & CEO, VibeLogic
  3. Aug 12

    Will the AI Data Center Backlash Really Make a Difference?

    For three years, the AI infrastructure story has been about chips, power, and capital. In 2026, a fourth variable arrived that the hyperscalers were not prepared for: organized, well-funded, and increasingly successful local opposition to data center development. Ray and Peter walk through the numbers behind the fight. Data Center Watch tracking shows at least $130 billion in US projects blocked or delayed in the first quarter of this year alone, with Carbon Direct putting the cumulative figure at $170 billion since 2024. Set against roughly $725 billion in projected 2026 capex across Alphabet, Amazon, Meta, and Microsoft, the question for operators is whether this is a genuine constraint or noise around the edges of an inevitable buildout. The bigger shift is who carries the risk. With digital infrastructure funds raising $157.6 billion in 2025 according to PitchBook, and private capital taking positions like Blue Owl's 80 percent interest in Meta's $27 billion Hyperion campus, a permitting fight in a single county is now underwriting risk for pension funds and insurers nationwide. Delay no longer just moves a launch date. It extends the payback period and compresses return on invested capital. In this episode: The scale of the buildout: 12 gigawatts of national compute capacity today, with the industry targeting 60 gigawatts by the end of the decade Who is writing the checks: OpenAI's $1.1 trillion in total infrastructure commitments, Anthropic's roughly $350 billion in compute commitments, and the rise of third-party capital as the load-bearing wall Why projects slip: 75 percent of capacity under construction already pre-leased, 1.6 percent vacancy, five-year grid interconnection backlogs, and 3-5 year transformer lead times What is driving residents into council meetings: $29.4 billion in added PJM customer costs, utility bills up as much as 267 percent in some markets, Google's water consumption up 34 percent year over year, and a Gallup finding that 71 percent of Americans oppose a data center in their own community, a higher share than opposes living near a nuclear plant How communities win: 833 active opposition groups across 49 states, more than 300 municipal bans and moratoriums since 2023, and Monterey Park's permanent ban approved with 88.34 percent of the vote Fifty different rulebooks: New York's statewide permitting pause, Ohio's 85 percent capacity payment requirement, Virginia's energy consumption tax, and West Virginia moving in the opposite direction The three playbooks that get projects built: Meta's community investment model in Louisiana, the legal route in Michigan, and Microsoft's quieter university and water reuse partnership in Wisconsin The operator takeaway: The buildout is getting slower and more negotiated, not stopped. For anyone modeling enterprise AI unit economics, the bottleneck has moved from chips to concrete, copper, and zoning boards. Cost models built on an assumption of falling compute prices need a second look this quarter. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

    Will the AI Data Center Backlash Really Make a Difference?
  4. Aug 4

    Forward-Deployed Engineers (FDEs) - AI's New ROI Battleground

    AI spending continues to accelerate, but the ROI story has not kept pace. This week, Ray Rike and Peter Buchanan dig into the proposed fix that has the market talking: turning frontier AI labs into professional services firms through the forward-deployed engineer (FDE) model. The role is not new. Palantir built the FDE function in 2005 to embed technical teams directly within customer environments, and by 2016, it had more FDEs than software engineers. What is new is the capital. Five ventures from the major AI labs and top three hyperscalers have committed more than $10 billion, betting that the bottleneck to enterprise AI value is not the model; it is getting that model wired into a customer's data, processes, workflows, and compliance requirements. Ray and Peter connect this moment back to the ERP era, when SAP and Oracle needed four to five dollars of services for every dollar of software, and explain why the same people, process, and services reality is playing out again with agentic AI. What the episode covers: Why OpenAI's $4 billion DeployCo, with a guaranteed 17.5% investor return, is the most aggressive and most financially puzzling bet of the fiveHow Anthropic (Ode), AWS, Microsoft, and Google each took a different structural path, from joint ventures to capital-light partner ecosystem playsWhy the MIT 95% pilot failure stat and McKinsey's finding that two-thirds of organizations have not started scaling AI make this a real problem, not a fringe oneHow incumbent consulting firms are playing every side at once to protect their AI practicesAlex Karp's argument that the AI industry broke its own business model, and the irony of him making it What CFOs and GTM leaders should take away: Ask any FDE partner for two or three production use cases with measurable outcomes before expanding scopeStart narrow with one win, not four or five simultaneous projectsPrice for outcomes up front, before mid-project renegotiationModel the ongoing maintenance cost, since 20 to 40% of the initial investment often goes to keeping it running, and an outside FDE team at $300 to $800 an hour is an expensive long-term maintenance line itemWeigh the neutrality trade-off honestly, since lab-backed ventures get you the deepest model roadmap access but also the deepest lock-inLawyer up on data privacy and IP, because these engagements can feed your proprietary workflows back into the next model release Ray's read: independent consulting firms are the likely long-term winners, with hyperscalers close behind. The AI labs are, in Peter's words, still leaving the Shire without their full posse ready to go. Listen to the full breakdown, and for the deeper analysis behind each big story, subscribe to the AI to ROI newsletter at ai2roi.substack.com. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

    Forward-Deployed Engineers (FDEs) - AI's New ROI Battleground
  5. Jul 30

    Chinese Open Weight Models Overtake US Frontier AI: What Every Enterprise Executive Needs to Know

    Twelve months ago, US frontier models controlled roughly 70 percent of AI traffic. Today, Chinese open-weight providers, led by DeepSeek, Z.ai, Moonshot, and Minimax, account for 45 to 61 percent of top-tier model traffic on OpenRouter, with DeepSeek alone processing more tokens than Google, Anthropic, or OpenAI individually. In this Big Story edition, Ray Rike and Peter Buchanan unpack how this shift happened, how US labs and regulators are responding, and three scenarios for how the closed versus open weight competition plays out for enterprise AI buyers. Key topics discussed: The pricing collapse driving enterprise migration. DeepSeek made a 75 percent price cut permanent in May, bringing its V4 Pro model to a fraction of a cent per million tokens versus $2.50 per million for GPT 5. Minimax delivers GPT 5.5 class coding performance at 5 to 10 percent of the cost. This is why Uber, Microsoft, and Walmart are now implementing formal usage governance on frontier models rather than treating cost control as temporary. The Mythos and Fable shutdown as a trust event. The 18-day suspension of Anthropic's top models over export control concerns spooked global enterprise buyers who realized mission-critical workloads could be cut off without warning. This single event accelerated the adoption of open-weight alternatives and pushed allied governments to invest in sovereign AI capacity. Distillation attacks and the IP leakage problem. Anthropic accused Alibaba's Qwen lab of running a large-scale adversarial distillation campaign, using tens of thousands of accounts and tens of millions of exchanges to extract agentic reasoning capability from Claude. This reframes the security conversation from model safety to unauthorized technology transfer, which is a distinct and arguably bigger risk for any enterprise relying on proprietary model capability as a moat. Cybersecurity parity is closing faster than expected. Multiple Asian labs, including Z.ai's GLM 5.2, Beijing based 360 Security, and Japan's Sakana AI, now claim benchmark performance approaching Anthropic's Mythos model on vulnerability detection and both offensive and defensive cyber tasks, often at significantly lower compute cost. This weakens the safety and capability gap argument that has justified restricting access to frontier models. Real deployments have moved from theory to production. Coinbase cut AI spend in half after migrating to Z.ai and Moonshot's Kimi models, even as token usage grew. Cursor shipped a coding tool built on Kimi, with a Grok-based version reportedly imminent. Andreessen Horowitz estimates 80 percent of its portfolio companies already use open-weight models in production AI products. Three scenarios for how this settles, and why the decision belongs at the board level. The hosts outline bifurcation (premium closed models for regulated use cases, open weight for commodity workloads), export control entrenchment (Washington treats the Fable ban as a template rather than a one-off), and capability convergence (the rationale for unilateral bans erodes as the performance gap closes). All three are already visible simultaneously, which means enterprise AI architecture decisions, including primary and backup model orchestration, are becoming strategic decisions that belong with the CEO and board, not just the technical team. Why does this podcast episode matter for enterprise executives selecting models, especially for agentic AI deployments? The vendor you choose today may not be the vendor you can use tomorrow, for reasons that have nothing to do with model quality. Regulatory risk, geopolitical exposure, and pricing volatility are now first-order variables in model selection, alongside capability and cost. Any agentic AI architecture built on a single model provider carries concentration risk that didn't exist a year ago. Building orchestration flexibility with primary and backup models, and understanding the true economics behind token pricing, is quickly becoming a board-level governance question rather than a procurement detail. For the metrics and framework to instrument and report ROI on your AI investment, get the Big Book of AI Metrics at benchmarkit.ai under Media. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

    Chinese Open Weight Models Overtake US Frontier AI: What Every Enterprise Executive Needs to Know
  6. Jul 28

    AI is Hot - AI Regulation is a Hot Mess

    On June 12th, the US Department of Commerce ordered Anthropic to suspend global access to Fable 5 and Mythos 5, its most powerful frontier models, with no advance warning and criminal penalties attached. In this week's Big Story episode, Ray Rike and Peter Buchanan walk through the first use of export controls to take a deployed frontier model offline, why it backfired for security rather than strengthening it, and what a coherent federal AI regulatory framework would actually need to look like. The shutdown mechanics. Because Anthropic could not verify user nationality in real time, the directive knocked out access for every user globally, including Anthropic's own non US employees and more than 150 companies across 15 countries running critical infrastructure workloads. Three reaction tracks. Industry, the developer community, and allied governments each responded differently. OpenAI pushed back on talent restrictions while its legal team blocked coordination with Anthropic on antitrust grounds, and a coalition of 150 cybersecurity leaders published an open letter asking not for less regulation but for a transparent, science based process. The regulatory vacuum, in five forces. Ray and Peter unpack the drivers behind what they call a hair on fire crisis: capability jumps outpacing any statutory framework, a state level policy tsunami of over 1,500 AI bills across 45 states, data center community opposition, the unresolved Anthropic and DOD dispute, and a bipartisan Congressional letter questioning why comparable models were treated differently. A six part framework. The episode lays out proposed solutions modeled on existing regulatory precedent: global market access agreements similar to military sales processes, mandatory pre release testing run by NIST modeled on FAA certification, clear federal versus state jurisdictional lines modeled on pharmaceutical regulation, expanded child safety authority, FERC fast tracking for data center grid access, and coordinated environmental standards. What enterprise leaders should do now. Ray closes with practical guidance: review AI vendor contracts for shutdown protection since force majeure clauses were never written with export controls in mind, build fallback infrastructure for critical AI dependent workflows, delay rushing into brand new model releases, and automate tracking of the fast growing state regulatory landscape. Full details are in the AI to ROI newsletter at ai2roi.substack.com. Subscribe, and consider reaching out to your representatives on this one. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

    AI is Hot - AI Regulation is a  Hot Mess
  7. Jul 22

    Building the 100x Org - The CFO as AI Architect with Dan Zhang, CFO & CBO at ClickUp

    Most companies treat AI as a layer they add on top of how work already gets done. Dan Zhang, Chief Business Officer and CFO at ClickUp, argues that it is exactly backward. In this episode, Ray sits down with Dan to unpack the "100x Org," ClickUp's framework for rebuilding the business around AI rather than sprinkling tools and tokens on top of a human-driven workflow, and why that distinction determines whether an AI initiative shows up as activity or as income statement impact. The conversation covers: Why ClickUp expanded the CFO's charter to own AI transformation end to end, after both a top-down mandate and a bottoms-up experimentation push failed to produce results that made it into production The "jobs to be done" framework Dan uses to separate primary work that actually moves the business from secondary work that just generates busy AI activity, and why most companies have a work redesign problem before they have an AI problem Dan's psychological test for AI ROI (would you pay for it with your own money) and why ARR per headcount is the right metric, but a lagging one that plays out over years, not weeks How ClickUp instrumented daily, not monthly, visibility into AI cost and token consumption, and why over half of enterprise companies in Ray's own research have blown their AI budget by 25% or more Why the real anxiety CFOs have about AI ROI isn't the return, it's confidence in cost management, and how a governance layer turns that anxiety into a guardrail instead of a monthly surprise Dan's rapid-fire advice on who should own AI ROI measurement, the two-to-three variables every CFO needs in place, and how early and mid-career professionals protect their relevance by owning the question, not just the answer If your organization has AI activity but can't yet point to AI impact, or your CFO isn't sure whether AI spend is a competitive advantage or a leak, this episode gives you the operating model and the financial discipline to tell the difference. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

    Building the 100x Org - The CFO as AI Architect with Dan Zhang, CFO & CBO at ClickUp
  8. Jul 21

    AI is a Compensation Scale Expense

    Token prices have fallen 98 percent since GPT-4, but enterprise AI bills are up 320 percent. In this week's Big Story episode, Ray Rike and Peter Buchanan trace where that money is actually coming from, and the answer is not the software budget. Drawing on Gartner, Oxford Economics, Zylo, Challenger Gray and Christmas, and Goldman Sachs data, the two lay out why labor, not IT, is becoming the primary funding source for AI at scale, and why almost no company has the measurement infrastructure to manage it. The price paradox. Per token costs have collapsed, but usage has grown faster than costs have fallen. Ray walks through the math behind average enterprise AI budgets rising from $1.2 million to $7 million in two years. Three cautionary tales. Uber consumed its entire annual Claude Code budget in under four months, Microsoft revoked thousands of Claude Code licenses over cost, and one unnamed enterprise ran up a $500 million bill in a single month. Ray and Peter break down why each was a governance failure rather than a technology failure. Only two budget pools are big enough. The IT and software budget represents just 3 to 4 percent of revenue, while labor represents 25 to 40 percent depending on industry. Ray makes the case that labor is the only pool large enough to absorb the AI spending trajectory Gartner and Oxford Economics are projecting. The attrition lever. Ray and Peter unpack how not backfilling open roles has quietly become the primary way enterprises are funding AI investment, supported by data showing over 113,000 tech layoffs in 2026 with 48 percent explicitly attributed to AI. Revenue per FTE as the tell. Ray shares benchmark data showing SaaS company revenue per employee up 25 to 35 percent over the last twelve quarters, and explains why this metric will be the clearest signal of whether the AI budget transfer is actually working. Five metrics every CFO needs now. The episode closes with a practical starting list: AI spend as a percent of revenue, AI spend per employee, inference spend as a percent of opex, inference cost as a percent of COGS for AI enabled products, and revenue per FTE tracked against labor cost and agent cost as a percent of OPEX. AI to ROI is always looking for guests with real-world examples of measuring AI budget impact. Reach out Ray on LinkedIn (@rayrike). Subscribe to AI to ROI at ai2roi.substack.com for the full June 9th edition, and leave a review wherever you listen See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

    AI is a Compensation Scale Expense
4.8
out of 5
41 Ratings

About

AI to ROI is a podcast that shares how enterprises translate AI investments into measurable business value. Hosted by Ray Rike, Founder and CEO of Benchmarkit, the show features senior enterprise leaders and AI software executives who share how AI initiatives move from pilots to production, and how ROI is actually measured and achieved. In addition, each week, we publish a bonus episode with AI to ROI Newsletter co-author, Peter Buchanan to discuss the Big Story of the Week. The AI to ROI podcast is the evolution of the original "Metrics to Measure Up" podcast.

You Might Also Like