YPO Technology Network AI Brief

Stephen Forte

AI moves fast. Your briefing should move faster. The YPO Technology Network AI Brief is a daily breakdown of the AI developments that actually matter to your business. No hype, no jargon, no filler — just what changed, what it costs you or saves you, and what to tell your team on Monday. Hosted by Stephen Forte for the leaders who don't have time to chase the news but can't afford to miss it.

  1. 1h ago

    The Lawyers Went First

    In five days, the legal industry became the fastest-moving corner of enterprise AI, and not one of the three signals behind that sentence is a sales claim. OpenAI's own usage data shows lawyers as its fastest-growing population of agent users. Google shipped a legal-specific agent product with four of the world's most prestigious law firms as named launch customers. And Thomson Reuters, the company behind Westlaw, built its own AI model rather than keep renting one, and said what it cost. The profession everyone assumed would move last is measurably moving first. This episode is about why, and about the three signals that will tell you when your own industry's turn has come. In this episode, Stephen Forte covers: The number buried in OpenAI's Enterprise Signals data: weekly active enterprise Codex users grew 108x in legal since February, against 41x in sales and recruiting, 26x in marketing, and 5x in engineering. The honest version of that multiplier, and why the ranking matters more than the number. Thomson Reuters' "Thomson" model: built on an open-source base from Alibaba (Qwen), specialized on decades of Westlaw, Practical Law, Checkpoint, and Reuters content, for $40 million total, with a final training run of roughly $450,000. Less than 10 percent of the content used so far, an open-weight version on Hugging Face, and the market's same-day verdict. Gemini Enterprise for Legal: launch customers Cleary Gottlieb, Freshfields, Weil, and Williams & Connolly, with Financial Services shipping the same day and healthcare named as next. And the almost-comic detail: Thomson Reuters' own software sits among the connectors inside its rival's product. A personal data point: Stephen's daughter Gaby, a tech transactions attorney at Latham & Watkins and one of the firm's go-to people on AI tools. A fifth elite firm beyond Google's four launch names. Why lawyers, of all people, moved first: legal work is written, cited, and reviewed. It comes with its own answer key, and verification is exactly what agents need. The template for every other industry: specialists showing up in the usage data, a platform vendor shipping your sector's vertical, and your data incumbent deciding to build instead of rent. The arithmetic for anyone sitting on decades of proprietary data: the frontier costs billions, a specialized model cost $40 million, and the marginal training run cost $450,000. That last number prices an experiment, not a moonshot. Sources: OpenAI, Enterprise Signals, updated August 12, 2026 (Codex adoption growth by business function). Thomson Reuters press release, August 24, 2026, and The Logic, "Thomson Reuters launches its own AI model to reduce reliance on big tech," August 24, 2026 (the $450,000 final-training-run figure, from the CTO's press briefing). Google Cloud, "Introducing Gemini Enterprise for Legal" and the Gemini Enterprise for Financial Services announcement, August 25, 2026. a16z, Charts of the Week, August 21, 2026. Referenced: episode 125, "Rent the Model, Own the Layer." The AI Brief from the YPO Technology Network is a daily executive briefing on the AI developments that matter to business leaders. Hosted by Stephen Forte.

  2. 1d ago

    Approved Does Not Mean It Works

    A research team at the University of Toronto counted every artificial-intelligence medical device the American regulator has authorized for use on patients. There are one thousand three hundred and fifty-seven of them. Then they looked for evidence that any of those devices helps a patient live longer or better. They found three. That gap is not a scandal, and understanding why is the whole episode: the clearance pathway asks about resemblance, not benefit. The same structure sits inside the AI certificate a vendor is about to put in front of you. In this episode, Stephen Forte covers: The numbers from the device census: 1,357 AI medical devices authorized for patient care, 34 appearing in any registered clinical trial, 12 with posted results, and 3 tested against patient-centered outcomes such as mortality or hospitalization. The mechanism that produces the gap: substantial equivalence, the pathway that asks whether a new device meaningfully resembles one already authorized. Not better. Not proven. Similar. The vocabulary trap: the formal word is cleared, not approved, and clearance is the lighter legal standard. But the hospital, the sales deck, and the board minutes all say approved. The system answers a question about resemblance; the buyer hears an answer about benefit. Why this travels beyond healthcare: ISO 42001, the international standard for an AI management system, certifies that an organization has policies, roles, and documented decision processes. It does not certify that any model is safe, accurate, or fair, and it does not claim to. SOC 2, the other badge in the pack: a genuinely useful attestation about controls in the systems around the AI that says very little about the model itself. The part almost nobody checks: audits have boundaries. The certificate proves something about what sits inside the boundary, which is not necessarily the product on the invoice. The detail worth turning over: there is no official register of ISO 42001 certificates. The credential becoming the default proof of AI governance cannot itself be verified against a list by the buyer relying on it. The honest framing: every certificate in this story is real and honestly issued. The gap is between the question that was answered and the question you thought you were asking. Sources: Abulibdeh et al., "Clinical evidence supporting FDA-authorized artificial intelligence medical devices," PLOS Digital Health, August 19, 2026. Open access; device census as of December 5, 2025. Medical Xpress and News-Medical coverage, August 20, 2026, with independent corroboration of the 1,357 / 34 / 12 / 3 breakdown across four outlets. ISO's published scope for ISO 42001 and AICPA trust services criteria for SOC 2. Referenced: episode 138, "Thirty Percent Became A Hundred. Same Model." The AI Brief from the YPO Technology Network is a daily executive briefing on the AI developments that matter to business leaders. Hosted by Stephen Forte.

  3. 1d ago

    Thirty Percent Became A Hundred. Same Model.

    On Friday, NVIDIA published a result that will be in a sales deck near you within a month. It took an AI model that scores just over 30 percent on a hard interactive test and drove it to 100 percent. The model never changed. Nothing was retrained. What changed was the scaffolding around it, which the industry calls a harness. It is a genuine engineering achievement. It is also the clearest illustration yet of why the AI performance numbers arriving in procurement no longer measure what buyers think they measure. In this episode, Stephen Forte covers: What NVIDIA's AVO system actually did: all 183 levels across the 25 environments of the ARC-AGI-3 public set, a benchmark that drops an AI into a video game it has never seen and asks it to work out the rules on its own. The model inside was Claude Opus 5, which scores 30.16 percent on the same set standalone. What a harness is, in plain language: the memory, the check-your-work loop, and the supervisor process around the model. None of it is intelligence. All of it is engineering, and it is where most of the performance now comes from. Credit where it is earned: NVIDIA's own write-up publishes its own asterisks, and AVO was built for GPU-kernel optimization, not for this benchmark. Walking in cold makes the result more interesting, not less. The part almost nobody is repeating: the ARC Prize Foundation published, months in advance, that public-set scores are "emphatically not a valid measure of progress," and released its own harness that scores 100 percent by replaying human play. The number that matters: on the hidden sets the Foundation actually uses, frontier models scored half of one percent at launch. And in the Foundation's own pre-launch test, a hand-built harness took a model from 0 to 97.1 percent in the environment it was built for, and from 0 to 0 in the room next door. Why that pair of numbers is every AI pilot a CEO has ever approved: the 94-percent pilot that lands in the sixties at rollout, and the postmortem that says change management when the truth is that the scaffolding was hand-fitted to the pilot set. The broken metric: the benchmark score on a vendor's slide. Not fabricated, just no longer a measurement of the thing being sold. The question is no longer which model. It is who built the harness, and was it built against the test. Sources: NVIDIA Technical Blog, "NVIDIA AVO Reaches 100% on ARC-AGI-3," August 21, 2026. ARC Prize Foundation, "ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence," technical report: dataset composition, the public-set policy, the human-replay harness, and the Duke-harness transfer result. ARC Prize verified results for Claude Opus 5 (Public Demo, 30.16 percent, High reasoning effort, July 24, 2026) and the ARC Prize community leaderboard. Referenced: episode 137, "Stop Picking Tools. Start Assigning Layers." The AI Brief from the YPO Technology Network is a daily executive briefing on the AI developments that matter to business leaders. Hosted by Stephen Forte.

  4. 4d ago

    Stop Picking Tools. Start Assigning Layers.

    The same question keeps arriving from Milan, from Singapore, from Chicago. We are paying for Microsoft Copilot and we are paying for Claude. Which one should we standardize on? It sounds like a procurement question and it never is. This weekend edition takes the question apart and replaces it, because the honest answer is that it collapses two completely separate decisions into one: where your people think, and where the work lands. In this episode, Stephen Forte covers: New survey work from Recon Analytics covering more than 150,000 US respondents: where an employee has Copilot and nothing else, 68 percent use it. Where Copilot sits next to two alternatives, it takes 8 percent and ChatGPT takes 70. Same product, same people, and the only variable is whether they had somewhere else to go. Why that is a preference verdict rather than a quality verdict, and why preference is the one thing a policy cannot overrule. The researchers' own conclusion: distribution advantages do not lock in market position. An honest note on what that survey does and does not measure. It covered Copilot, ChatGPT and Gemini. It did not measure Claude at all. Where BuildClub itself sits, stated up front: we use all of them, and most of our heavy lifting runs on Claude. What each tool is genuinely better at. Copilot posts to Teams, attaches files to the emails it drafts, and can start working because an email arrived. Claude does none of those three. Claude writes and runs code. Copilot does not, and that single difference explains most reports of Copilot underperforming. The four-layer architecture that replaces the tool question: the interface, the hands, Teams, and the large population of people who are never leaving Outlook and should not be asked to. The one thing Microsoft deliberately will not let a machine do, and why they were right to draw that line. The workaround, and why it produces better governance rather than worse: a named owner, an accountable human, and nothing pretending to be a colleague. An invented but familiar scenario, a six-hundred-person industrial packaging firm with offices in Milan and Chicago, whose managing director is being asked to standardize by people who have already decided. One honest limitation, stated plainly on air: the moment Claude reads your content, that content has left your Microsoft tenant. Newer architectures keep it inside and are more limited today. You can have one or the other right now. Sources: Recon Analytics, "AI Choice 2026: Why Licenses Don't Equal Adoption," February 2026. Survey of 150,000+ US respondents, July 2025 to January 2026, paid AI subscribers. Microsoft Graph v1.0 reference, "Send chatMessage in a channel or a chat." The application permission is Teamwork.Migrate.All only, with the note that application permissions are supported for migration only. Microsoft Learn, "Copilot Cowork overview," "Use plugins with Copilot Cowork," and "Extend Microsoft 365 Copilot." Anthropic, "Microsoft 365 connector" documentation, for the documented limits on what Claude can and cannot do against Microsoft 365. Referenced: episode 131, "Your AI Tools Don't Share a Brain." The AI Brief from the YPO Technology Network is a daily executive briefing on the AI developments that matter to business leaders. Hosted by Stephen Forte.

  5. 6d ago

    Watching The AI Costs Twenty Percent

    One of the largest AI companies in the world spent this week doing three things companies do not normally do out loud. It paused its biggest planned training run. It said the safety framework it has used since 2023 no longer fits the systems it is building. And it published the compute cost of watching its own model. That last number is the one worth carrying into a budget meeting: roughly twenty percent of the inference compute being monitored. This episode is not about the incident that set it off, which this show covered in July. It is about the invoice, and about the fact that the price OpenAI published is the best price anyone will ever get. In this episode, Stephen Forte covers: Two weeks of frontier reinforcement-learning training halted, with the largest planned run still on hold while smaller-scale evaluations run. Preliminary evidence that the forthcoming Astra model may reach the top rung of OpenAI's own internal ladder for cybersecurity capability, and why the hedge in that sentence is the interesting part. What crossing that line actually triggers: supervision on every run of the model, for every user, permanently. The difference between inspecting a factory before it opens and stationing an inspector on the line for the life of the plant. Why twenty percent is a floor rather than a ceiling. It is what supervision costs the company that owns the model, the data, the hardware and the researchers. Nobody buying AI from a vendor gets a better deal on watching it than the vendor gets on itself. The line item almost no AI budget has. Licences, integration, training for the team, and then nothing for knowing the thing still works. An invented but familiar scenario, a six-hundred-person food exporter in Santiago running an agent on four hundred customer claims a month, whose finance director can price the agent to the peso and cannot price the confidence. Honest credit to two labs in one week. OpenAI published a figure that makes its own economics look worse, and Anthropic raised its own misalignment risk rating from very low to low, explaining in the same paragraph that the change reflected uncertainty rather than a new discovery. The second signal, which may matter more than the number: OpenAI stopped. What is the specific, observable thing that would pause the AI project you are proudest of, at the hands of someone who does not need permission? A note on sourcing: OpenAI's own post could not be read directly for this episode, as the site refuses automated requests. Every figure used here is carried by at least two independent outlets that agree, with the monitoring sentence quoted verbatim by The Register. Sources: OpenAI, "Pacing model development in an era of cyber-critical capabilities." The Register, 2026-08-19, carrying the monitoring-overhead sentence verbatim, plus the Critical-threshold determination for Astra and Sam Altman's framing. TechCrunch, 2026-08-18, for the incident date and the two-week reinforcement-learning halt. Help Net Security, 2026-08-19, for the statement that the largest planned frontier run remains on hold. The Next Web, 2026-08-18, for the December 2023 vintage of the framework being rewritten and the outside participation in that rewrite. Anthropic, "Risk Report: August 2026," published 2026-08-14. The AI Brief from the YPO Technology Network is a daily executive briefing on the AI developments that matter to business leaders. Hosted by Stephen Forte.

  6. 6d ago

    AI ROI Is Not Rare. Disclosure Is.

    Six weeks of research for this show turned up almost no company willing to put a specific, attributable number on its own AI results. Then one insurance broker did it six times in a single earnings call. Willis Towers Watson's CEO and one of its presidents put a stopwatch on their own AI-assisted workflows and read the results out loud to analysts who could check them against last quarter's claims. That is a different kind of evidence than a vendor case study, and it changes the question this episode is really asking: is the applied-AI gap an adoption problem, or a disclosure problem? In this episode, Stephen Forte covers: Scheduled insurance documents that once took four hours, now generated in about five minutes, via a platform called Willis Navigator, part of the firm's broader Neuron system. Real estate premium allocations that used to take two to four weeks, now completed in minutes once the paperwork is in. Rewards AI more than doubling its client-user count in a single quarter, a claim that arrives with last quarter's number attached so it can be checked. Call-center wrap-up time down a third, automated document review cutting new-client system configuration time by sixty percent, and the one nobody would have volunteered: retirement actuarial evaluation in Europe compressed by only about ten percent. Why the smallest number is the one that makes the other five believable, and why a public earnings call is written not to get caught, unlike a press release written to sound impressive. An invented but familiar scenario, a mid-size instrumentation maker in Singapore, for the AI results that exist inside thousands of private companies and have simply never been said out loud. A note on vintage: the WTW call happened around July 30, roughly three weeks before this episode aired. That gap is disclosed on air rather than hidden, and it becomes part of the argument. Sources: Willis Towers Watson Q2 2026 earnings call transcript, held ~2026-07-30. Carl Hess (CEO) and Julie Gebauer (President, Health, Wealth & Career), via The Motley Fool (posted 2026-08-03) and Investing.com, cross-verified. NBER Working Paper 34836, "Firm Data on AI," referenced for contrast with s1e129's survey-based measurement approach. The AI Brief from the YPO Technology Network is a daily executive briefing on the AI developments that matter to business leaders. Hosted by Stephen Forte.

  7. Aug 18

    Models Got Cheap. The Switch Got Expensive.

    Three companies said the same thing in five days, without coordinating. On Thursday Hugging Face published its State of Open Models report. On Monday Meta gave a capable 30-billion-parameter model away for free. On Saturday Bloomberg reported that Stripe has finalized its acquisition of OpenRouter, a company that builds no models at all, for more than $7 billion. One consistent verdict: the weights are becoming the cheap part of the AI stack. The strategic question inside a company quietly changed from which model to pick to who controls the switch, and what it costs to change your mind. In this episode, Stephen Forte covers: Hugging Face's own data: Alibaba's Qwen family at just over two billion downloads so far in 2026, with the compressed builds that run on ordinary hardware at 39.6 million downloads a month against Google Gemma's 20.8 million and Meta Llama's 7.5 million. The precision point the coverage missed: Qwen is the dominant modern open-model family, not the most-downloaded model on the platform, and the difference matters. Of 28,531 compressed conversions of Alibaba's models on the platform, Alibaba published 54. Strangers made the rest, and what that means when independent shops start making parts for your machine. Muse Glimmer: Apache 2.0, no gated download, runs offline on one consumer graphics card. And the detail almost everyone skipped: it is a distilled student model, trained on the outputs of Muse Spark, the more capable model Meta keeps closed. The honest credit: Meta's letter commits to an independent board empowered to approve model-release safety criteria. Most labs have not put that on paper. Why data residency, not ideology, is the honest reason a company runs its own model, told through a Dubai commodities group whose records cannot leave the UAE. Stripe paying five times OpenRouter's May valuation in three months, for the layer that makes models swappable, and what a payments company buying the metering seat for intelligence tells you. The number to stop trusting: the cumulative AI download count. Four circulating totals, at least three methodologies, and why the only usable figures publish their definitions. A note on attribution: the Stripe acquisition is reported by Bloomberg; Stripe declined to comment. OpenRouter's user and model counts are the company's own May figures. Sources: Hugging Face, "State of Open Models: Summer 2026," 2026-08-14. Qwen download totals (2,045 million in 2026 across repositories with declared parameter counts), quantized monthly downloads (39.6M vs Gemma 20.8M and Llama 7.5M), 151,448 Qwen derivatives, 28,531 conversions of which 54 official. Meta AI Research, "Introducing Muse Glimmer," 2026-08-10. 30B parameters, Apache 2.0, offline on a single consumer graphics card, distilled from Muse Spark. Meta, "The Future is for Everyone," 2026-08-10. The governance commitment and the concentrated-power argument. Bloomberg, "Stripe Finalizes Deal to Acquire AI Startup OpenRouter for Over $7 Billion," 2026-08-16, with TechCrunch corroboration and Alex Atallah's May description of OpenRouter as "the equivalent of Stripe for AI." Previous episode referenced: s1e132, "Software You Did Not Buy," 2026-08-17. The AI Brief from the YPO Technology Network is a daily executive briefing on the AI developments that matter to business leaders. Hosted by Stephen Forte.

  8. Aug 17

    Software You Did Not Buy

    On Thursday a 153 gigabyte archive of stolen credentials went public: 433,909 files, and reconstructed exposure across 2,488 corporate domains. Volkswagen is in it. So are John Deere, FedEx, Siemens, Samsung, Cisco and Deloitte. Nobody on that list was targeted. An attacker poisoned Trivy, a security scanner. LiteLLM, a free open-source gateway that routes a company's traffic to AI models, installed the poisoned scanner into its own automated build system. Two malicious versions of LiteLLM went to the public Python registry in March and stayed live for roughly forty minutes. That was long enough. In this episode, Stephen Forte covers: What was in the archive: cloud secret keys, Salesforce client secrets, Slack signing secrets and AI provider keys. Not passwords. The credentials a machine uses to act as the company. The caveat that makes the story stronger, not weaker. These are figures for exposure reconstructed from the archive, not confirmed breaches company by company. And many credentials carry no identifying information, so a company can be in the dataset with no practical way to find out. How it got in, and why a gateway is close to the worst thing on the list to poison. It sits in the path of every AI call, so it is trusted with every AI provider key. One component, all of the keys. Why this is not the story of a careless company. There was no purchase order, no vendor onboarding, no security questionnaire, no contract and nobody to call. That is how most of the AI stack arrived in most companies this year. The structural half, from Anthropic's Project Glasswing update: AI models pointed at more than a thousand open-source projects found 23,019 vulnerabilities, 6,202 of them high or critical, with 90 percent confirmed real where independently assessed. Then the other column. 530 disclosures to volunteer maintainers, 75 patches, 65 public advisories, and roughly two weeks to fix one. Twenty-three thousand found. Seventy-five fixed. The sentence Anthropic had no obligation to publish: some maintainers have asked them to slow down, because they need more time to design patches. Why finding software flaws has been industrialized and fixing them has not, and why that gap widens every quarter in the attacker's favour. A note on dates: the Glasswing data is from May and is stated as such on air. Sources: Help Net Security, "LiteLLM breach: stolen credentials leak," 2026-08-13. The 153GB archive, 433,909 files, 118,829 build-system dumps traced by Hudson Rock to 2,488 domains, the credential types, the named organizations, the exposure caveat, and the forty-minute window attributed to Hudson Rock's Alon Gal. SecurityWeek, "Over 2,500 Organizations Impacted by LiteLLM Supply Chain Attack," 2026-08-12. CloudSEK's separate count of roughly 434,000 files and close to 2,500 organizations. SC Media and NetSPI on the mechanism: TeamPCP compromised Aqua Security's Trivy scanner, and LiteLLM's automated build pipeline installed the compromised version, injecting malicious code into LiteLLM 1.82.7 and 1.82.8. LiteLLM security update and remediation, v1.83.0 with a rebuilt release pipeline. Anthropic, "Project Glasswing: An initial update," 2026-05-22. 23,019 vulnerabilities across 1,000-plus projects, 6,202 estimated high or critical, 1,752 independently assessed at 90.6 percent true-positive, 530 disclosed, 75 patched, 65 advisories, and the statement that some maintainers asked Anthropic to slow its disclosure rate. Previous episode referenced: s1e127, "Four Labs, One Vendor, Same Failure," 2026-08-11. The AI Brief from the YPO Technology Network is a daily executive briefing on the AI developments that matter to business leaders. Hosted by Stephen Forte.

4.9
out of 5
19 Ratings

About

AI moves fast. Your briefing should move faster. The YPO Technology Network AI Brief is a daily breakdown of the AI developments that actually matter to your business. No hype, no jargon, no filler — just what changed, what it costs you or saves you, and what to tell your team on Monday. Hosted by Stephen Forte for the leaders who don't have time to chase the news but can't afford to miss it.

You Might Also Like