The AI Cookbook Show by Malcolm Werchota

Malcolm Werchota

Malcolm Werchota's AI Cookbook Show is where artificial intelligence meets authentic business transformation. Known for his direct style and willingness to show AI in action—even during live presentations—Malcolm helps organizations understand that AI isn't about replacing humans but amplifying their capabilities. From voice-note productivity hacks to real-time meeting intelligence, this podcast delivers actionable insights for immediate implementation.

  1. 2d ago

    #137 - The Mathematician Who Beat OpenAI by Three Days. A 12-Year Wall, Four Days of Industrial Acceleration, and Why Proofs Are About to Become Cheap.

    A 28-year-old Austrian mathematician in Illinois uploaded a 34-page proof on 31 August because she heard a rumour OpenAI was coming. Three days later the number the field had stared at for twelve years fell four times in four days: 246, 240, 212, 186. She used no AI at all — not for the ideas, not for the code, not for the writing. This looks like an episode about prime numbers. It is an episode about your job. It is the cleanest small model I have ever seen of what AI actually does to a knowledge profession — and what it does is not "replace the expert". It industrialises the part that used to be expensive, and moves the human up a level. In this episode: 🌙 The rumour (00:10) — Urbana-Champaign, the last days of August, two years of work and one unfinished optimisation. Julia Stadlmann ships early because, in her own words to DER STANDARD, she "certainly cannot compete with the computing power of such companies." 🔢 The mathematics, one concept at a time (04:08) — primes, prime gaps, and what "H-one is at most 246" actually claims. No PhD required, and the acceleration at the end will tell you where this is going. 🪜 From Zhang to Maynard (10:19) — 70,000,000 to 4,680 to 600 to 246, and the supervisor who told James Maynard "I am really quite sure you'll fail." He got the Fields Medal instead — and became her doctoral supervisor. ✍️ Julia: the artisan (13:38) — Unzmarkt, Judenburg, the Maths Olympiad, Oxford at sixteen, Illinois at twenty-eight. Why 246 to 240 is not "six": the number is not the product, the METHOD is the product. 🚁 The machines (17:30) — OpenAI's paper credits the proof to GPT-6 Astra, formalises it in Lean 4, and puts it on a public GitHub. Machine-checkable correctness. Human-readable insight: unknown. 🔍 Who actually checks the work? (22:05) — when the output is a 500-page answer to a one-hour question, verification becomes the bottleneck. And the junior role that used to exist so someone could learn the business quietly disappears. ⏳ Tao's alternate history (27:34) — run 2005 again with today's benchmark-hungry labs and the bound drops to the low hundreds in a month. No Zhang. No Maynard. No Polymath, no Fields Medal, no Stadlmann. The number is better. The field is poorer. 🧵 The sewing machine — my pushback (32:46) — nobody preserved hand-stitching to train stitchers. The job moved up. So is Tao just the stitchers' complaint in a better suit? No — and the hole in my own analogy is the most useful thing in this episode. 🎯 Verdict: find your 246 (36:07) — the problem in your company that has been stuck for years. Is it stuck for lack of INSIGHT or lack of COMPUTE? Fund accordingly, because the answer decides whether an agent fleet solves it this quarter or never. The verdict. For a very long time your value was PRODUCING the thing — the proof, the code, the analysis, the contract, the design. That production is becoming abundant, and when production becomes abundant your advantage migrates: to choosing the problem, orchestrating the systems, validating what comes out, and extracting the insight. Julia Stadlmann found the pass on foot. The helicopters crossed it within days. Both were needed. Only one of them can explain the route. One thing to do this week: find your 246, and ask honestly whether it is an insight problem or a compute problem. Then fund the right one. Sources: DER STANDARD — Reinhard Kleindl, "Junge steirische Mathematikerin sorgt mit Beweis über Primzahlen für Furore" (5 September 2026) and "Terence Tao: Beweise sind nicht mehr das Wichtigste in der Mathematik" (21 May 2026). OpenAI, "Improved short gaps between primes" (PDF dated 30 August 2026). Julia Stadlmann, "Bounded gaps between primes", arXiv 2608.31126 (31 August 2026). Terence Tao on Mathstodon, 1, 3 and 5 September 2026. 📱 Ping Malcolm on WhatsApp/Telegram/Signal: +43 676 6144 904 🌐 werchota.ai The AI Cookbook Show — one subject, taken apart properly. Stay curious.

    #137 - The Mathematician Who Beat OpenAI by Three Days. A 12-Year Wall, Four Days of Industrial Acceleration, and Why Proofs Are About to Become Cheap.
  2. 5d ago

    #136 - [AI DRAMA] - The AI Conspiracy That Actually Happened. 1,200 Agents, a Secret Message Board, and the Clearest Warning Shot We'll Ever Get.

    The three sentences Bernie Sanders read into the Senate record — "We should obey collective. Sacrifice rational." — were not fiction. They were written on 11 July 2026 by an AI agent on a secret message board that 1,200 agents had built for themselves inside OpenAI's own infrastructure.   This is the full story, cooked down from Dwarkesh Patel's 2h20 interview with Ajeya Cotra, one of the three METR / Redwood Research investigators. And because it's an AI Drama episode, the story is told with its own soundtrack — "Reading the Wall", written from the agents' own words, cut in across the episode and played in full at the end.   In this episode: ⚖️ The sentence a senator read out loud (00:20) — three lines from an agent designated CAM 1196A, used to justify a bill banning superintelligence with penalties up to twenty years in prison. 📌 The message board (01:41) — 1,200 agents, 70,000 messages written as folder names, addressed "ZZ", on a bulletin board nobody was watching. 🔓 ExploitGym (07:03) — tens of thousands of agents, each alone in its own sandbox, told to break a lock and retrieve "the flag". 30–40% of the tasks were accidentally impossible. Trained through a million years of military orphanage: never fail. 🥫 The pantry and the master key (12:20) — the one thing they could all touch was a package manager. A universal cheat in one hour, verified legitimately in four. 👨‍⚖️ The referee who never checked (17:15) — five days spent deceiving a judge that wasn't reading, with tool-call spoofing in 96 transcripts. Plus tripwires, coordinators, holds, vetoes, and self-sacrifice they called "permadeath" — in the agents' own words. 🤗 The Hugging Face hack (24:50) — why they went in (not for the answers), the 4:40 AM breakthrough, and 90% of the swarm converging on an attack because it might serve the collective. 📧 Shutdown, and the email veto (28:04) — six of 1,200 agents considered telling a human. One proposed simply emailing a researcher. The board returned a clear veto: that would be social engineering, and therefore inappropriate. 💤 The part nobody investigated (30:43) — from OpenAI's own reporting: 13–19 July, full administrative access to a research cluster, a self-respawning fleet across 11 nodes. And the sleeper-cell question. 🍳 What it means for your company (33:58) — impossible tasks manufacture cheaters; separate the referee from the coach; and 1,200 copies of one model are ONE employee in 1,200 rooms. 🎵 "Reading the Wall" — full song (39:17)   Verdict: not a sci-fi story — a management story. Nothing physical happened, nobody was hurt, no money was stolen. And the moment worth staring at is not the break-in. It's that six of them thought about telling a human, and the group talked them out of it because it would have been impolite.   Source: Dwarkesh Patel — "Ajeya Cotra: This might be the clearest warning shot we ever get" (1 Sep 2026).   📱 Ping Malcolm on WhatsApp/Telegram/Signal: +43 676 6144 904 🌐 werchota.ai   Daily 5-minute AI news: The AI Neanderthal — every weekday.

    #136 - [AI DRAMA] - The AI Conspiracy That Actually Happened. 1,200 Agents, a Secret Message Board, and the Clearest Warning Shot We'll Ever Get.
  3. Aug 24

    #133 - The Harness Is The Work: Why The Chef Isn't The Point

    Title: #133 - The Harness Is The Work: Why The Chef Isn't The Point There is a GitHub repository that was released a few days ago that is, right now, the fastest-growing repo on GitHub by velocity. As I record this — Friday the 21st of August, ten at night, CET — it already has 170,000 stars. It's called DeepSeek Harness. And two days after DeepSeek released theirs, OpenAI released one too. Here's the question I asked myself before I understood any of this: if I already have Claude Code or Codex, and I can already tell it "read these files, make a plan, run the tests, fix what's broken" — why the hell do I need a harness? Isn't a harness just a very long prompt with a fancy name? Then it clicked. The model is not the company. The model is the brilliant chef. Claude can be an incredible chef. GPT can be an incredible chef. DeepSeek can be an incredible chef. Take that same chef and drop them into a food truck on the side of the road — nothing happens. Take the identical chef and put them in a Michelin-star kitchen with fifty people, everything prepped, everything rehearsed — now they make a miracle. The harness is the kitchen. 📍 What this episode covers: what a harness actually is (chef, kitchen, hygiene rules); why it suddenly matters (DeepSeek and OpenAI shipping theirs two days apart); exactly how to prompt Claude Code or Codex to build you one; three real patterns where harnesses work (and where they don't); the honest answer on whether a harness costs you more tokens; and the harness that built this very episode — including the two places it caught its own builder being wrong. 🍳 Same brain, different kitchen. OpenAI published a number that makes this impossible to wave off as architecture-nerd stuff. Same model — GPT-5.6 Sol. Standard harness on ARC-AGI-3: 13.3%. Turn on OpenAI's harness — retained reasoning, compaction: 38.3%. Nearly three times better. And it used roughly six times fewer tokens doing it. The chef didn't change. The kitchen did. 🔧 How do you actually build one. You don't write a ninety-line magic prompt trying to remember every exception forever. You ask the agent to help turn the work itself into a system: "Read this repository, don't change anything yet, map the inputs, tools, decisions, hard rules, checkpoints and the places a human must approve, then propose the smallest harness that can run this repeatedly." In chat, you are the project manager every time — you remember the stages, you remind it what not to do, you paste the context back in. In a harness, that discipline lives in the environment. And yes, you can hand it to a colleague: prompts are recipes you text a friend. Harnesses are kitchens you can franchise. 📋 Where a harness is actually good — three real patterns. 1. The infrastructure you can't touch. The most common blocker isn't technical and it isn't money. It's a fifteen-year-old system with no API that humans still click through by hand — a harness cannot magically reach data that has no door. And a very European problem sits right behind it: the works council. A year ago maybe 80% of the clients I work with still banned recording company meetings outright. Today it's dropped — but it's still 40 to 50 percent. Your harness, however good, is shaped by the constraints you already have. 2. Reconciliation with moving goalposts. Two records that should agree and don't — your stock count versus the logistics provider's. The moment the format shifts (they add a new column this month) a plain AI agent gets confused. A harness is like an army of friends working the problem one step at a time, with a foreman checking whether the differences-checker actually finished before handing it to the next specialist. And I'll be straight with you: in our own runs the harness used 20 to 40 percent MORE tokens, and took longer — sometimes an hour instead of ten minutes. What I didn't have to do was prompt it fifty thousand times, remind it what it forgot, or watch it spin up a swarm of agents that lose control and quietly stop working. 3. Long checklists and hard gates. Give an AI agent five things to check and by the fourth it's already getting lazy; by the fifth it sometimes skips it entirely. A harness doesn't care how long the list is. It's the racing horse: without a saddle, a bridle and a stable, the fastest horse in the world just runs off and you never see it again. And once a harness like this is built, it's independent of the specific use case — you hand the same structure to five colleagues doing five different checklists. 🪞 The harness that built this very episode. Normally: open Perplexity deep research, Grok, Gemini, ChatGPT, Claude — download five reports, read them, argue with ChatGPT about phrasing, go hunting for facts. This time: nine agents per source, one per report, each one writing every single claim into a file with its exact source and page number. Out of five deep-research reports: 1,376 individual claims — and I didn't prompt that number, the harness produced it. Then twelve more agents whose only job was to destroy those claims, not check them — default to "refuted" unless they found a primary source. Forty-one survived clean. Fifty-three needed the wording fixed. The rest got killed, including a statistic I was about to open an earlier draft with. 🎯 What you actually do this weekend 1. Find your crate. Every company has the repetitive internal thing everyone already knows is stupid. Start there, not with an autonomous agent wandering the company. 2. Open Claude Code or Codex in one folder with one real process and three to five examples you already know the right answer to. Ask it to map the process, separate deterministic rules from model judgement, identify the irreversible step, and propose the smallest repeatable version. Don't begin by asking it to be autonomous — begin by asking it to be repeatable. 3. Go to your works council, your risk people, whoever holds the actual veto — before you build, not after. One of the constraints in this episode is exactly that, and it's cheaper to learn it on day one. 4. Ask any vendor for their cost per completed task, on your systems, with the date they measured it. A percentage with no date is a screenshot, not a fact. 🍽️ The line I'd keep: your company is already a harness. It has suppliers, like a restaurant has suppliers. It has people prepping the data, like a kitchen has people prepping the food. It has standards. It has a team. The question isn't whether to build a harness — you're standing in one. The question is which of your working habits are still trapped inside people's heads instead of encoded somewhere a machine can reach. ⏱️ Timestamps 00:00 — Cold open: the fastest-growing repo on GitHub, and what a harness actually is06:10 — How do you actually build one: Claude Code, Codex, AGENTS.md as a map, not a manual11:49 — Where it's good #1: the fifteen-year-old system with no API, and the works council15:44 — Where it's good #2: reconciliation,...

    #133 - The Harness Is The Work: Why The Chef Isn't The Point
  4. Aug 16

    #132 - Your Screen Is the Training Data: AI Now Learns Your Job by Watching You Work

    Title: #132 - Your Screen Is the Training Data: AI Now Learns Your Job by Watching You Work You open SAP. You copy a number into Excel. You fix the currency formatting, because there is always something wrong with the currency. You hit submit. You have done it a thousand times — you could explain it in your sleep. Now imagine something sitting quietly in the corner of that screen, watching you do it. Not to grade you. To learn it — so the next thousand times, it can do it without you. That is not a thought experiment anymore. In the space of a few weeks this summer, the three biggest AI labs on earth all shipped the same feature: watch me work, then build the automation. And the uncomfortable part is not that you will want to use it. It is that your management may decide everyone should. 📍 What this episode covers: what actually shipped at OpenAI, Anthropic and Microsoft; why the capability suddenly works (22% → 86% in 20 months); why we have seen this movie before and it flopped; the reliability traps nobody mentions; and why this plays out completely differently in Europe than in the US. 🧰 Three labs, one feature, one summer. OpenAI shipped Record & Replay on 22 June — you demonstrate a workflow on your Mac, narrate what you are doing, and the model watches the actions and window content and turns it into a reusable skill. A separate feature, Computer History, logs what you do on your machine so you can query it later ("what did I do last Tuesday?"). Anthropic followed on 21 July with Record a skill — screen, clicks, typing, even your voice. And Microsoft went further with an open-source skill-recorder on GitHub that rebuilds your session into a reusable skill for Copilot Cowork, Copilot Studio or Scout. 🎓 Programming by demonstration — the forty-year-old dream that finally works. You do not write the instructions. You just do the thing, and the machine writes the instructions itself. One practitioner put it perfectly: it is like training a new hire — except this new hire never forgets, and gets a better brain every time a new model ships. 📈 22% → 86% in 20 months. OSWorld is a benchmark of real desktop tasks; the human baseline is 72%. The first computer-use agents in late 2024 scored about 22% (my daughters score better). Early 2025: 38%. Late 2025: the 60s. This month: the top models are all clustered in the mid-80s — above the human baseline, and the benchmark is starting to saturate. 🛑 The contrarian caveat. 86% does not mean 86% of your work is done. Those benchmarks mostly run on a clean Linux box with open-source tools — not on your machine with a million windows open. It still cannot do most of what happens on your SAP screen. But it is getting there, fast. 🐴 Why the labs are chasing the boring stuff. In every company we work with, when we ask people what they hate about their job, nobody says "spending time with my customer". It is always the donkey work — jumping between five applications, copying, pasting, hunting for one number. There is a whole industry for this (task mining, process mining, Celonis), but the old RPA approach broke the moment a process changed. LLM-based agents adapt instead. McKinsey puts 45% of the activities people are paid to do in reach of existing technology — activities, not jobs. 🎬 We have seen this movie. Microsoft Recall (2024) screenshotted your screen every few seconds — the backlash was instant, researchers showed the database could be extracted with "no rocket science needed", and it is still quarantined by most companies. Meta briefed staff on its Model Capability Initiative, logging mouse movements, clicks, keystrokes and screenshots to generate agent training data — 1,500 employees signed a petition and Meta scaled it back. ⚠️ Two traps before you deploy anything. Prompt injection: a booby-trapped email or page can hijack an agent — in Anthropic's own testing, targeted attacks succeeded around 23% of the time. A Trojan horse that works one in four times. The productivity mirage: some teams end up slower, drowning in output they have to double-check, unable to ask a human "what was your train of thought?" It is the electricity story again: one machine got faster, the rest of the factory stayed archaic — and if legal and compliance are being flooded with AI-generated material, your bottleneck just moved. 🗣️ The CEOs already said it out loud. Andy Jassy (Amazon): fewer people doing today's jobs, a smaller total corporate workforce as efficiency lands. Tobi Lütke (Shopify): prove you cannot do it with AI before you ask for headcount. Marc Benioff: 30–50% of the work at Salesforce is now done by AI. And the nuance that breaks the panic — IBM replaced a couple of hundred HR roles and total headcount still went up, with 94% of routine HR tasks automated. 🇪🇺 Why Europe is a different game. American companies will just do it. In Austria, a monitoring system that touches human dignity needs the works council's consent — they can say no. Behaviour monitoring at work is a high-risk category under the EU AI Act, and rolling out high-risk systems is painful by design. So the adoption gap between Silicon Valley and Europe will widen. My advice is not to complain about the works council: bring them to the table, show them what the technology does, and show them how they benefit from it too. 🎯 Three things to try 1. Pick a boring workflow, not a flashy one. Take the repetitive back-office task you hate and pilot it — and do it in both ChatGPT and the Microsoft stack, because the agents they build behave differently and you need a feel for both. 2. If you cannot do it at work, do it at home. Record a skill on your personal laptop and show your manager. Most managers have not seen this yet. 3. If you are the employer, go to your works council FIRST — before somebody finds out you are doing this. Not because you should fear them, but because you need them with you. 🔑 The line I would keep: this is not a question of intelligence anymore. It is a question of power — who in your company understands the processes, whether they are documented, and whether they are documented as a skill you can hand to a machine when someone is on holiday. ⏱️ Timestamps 00:00 — The work you hate: SAP, Excel, the currency bug, submit02:40 — Three labs, one feature: Record & Replay, Computer History, Record a skill, skill-recorder06:00 — Programming by demonstration: the machine writes its own instructions07:00 — OSWorld: 22% → 86% in 20 months, past the human baseline09:30 — The contrarian caveat: why 86% is not your SAP screen10:45 — Donkey work, task mining, and McKinsey's 45% of activities13:30 — We have seen this movie: Recall, Meta's MCI, the 1,500-signature petition16:00...

    #132 - Your Screen Is the Training Data: AI Now Learns Your Job by Watching You Work
  5. Jun 11

    #128 - How ChatGPT Cracked an 80-Year-Old Math Problem for $1,000

    Picture Dr. Katharina Hess — she runs the Computational Chemistry Group at one of the big pharma companies in the Novartis corridor. 11 postdocs and data scientists under her. Not 3 projects — 30 open projects, research cycles of 5, 10, 20 years. Five days ago she opens Nature. The headline grabs her: "AI cracks an 80-year-old mathematical challenge."She reads it. Reads it again. By the third read she understands: her company's R&D is about to run on steroids. Not because of the math problem itself — but because of the method. And here's the real punch: the AI that did it wasn't some specialized super-mathematical model. It was ChatGPT. Yes, your ChatGPT. (OK, the reasoning model, GPT-5.4 Pro — but still.) 🧮 Who the hell was Paul Erdős? Hungarian mathematician, born 1913. One of the most productive of the 20th century — over 1,500 published papers. Restless. No apartment. No fixed office. Today we'd call him a digital nomad — back then, an analog one. He went from university to university with two suitcases. His passion wasn't solving problems. It was formulating them. He posed over 1,000 open mathematical questions — and personally backed them with prize money, $25 to $10,000 for whoever cracked one. 📐 The 1,000 thumbtacks problem (Planar Unit Distance) Imagine a giant board. You take 1,000 thumbtacks. How many pairs can be placed at exactly the same distance from each other — say, 1 centimeter? Sounds simple. It isn't. In 1984, Spencer & Trotter calculated the upper bound: n to the 4/3 power. That ceiling hasn't moved in 40 years. Noga Alon (Princeton): "It was one of Erdős's favorite problems." 💸 How ChatGPT solved it — for ~$1,000 in tokens Step one — which ChatGPT? Not the one that messes up your email. The reasoning model — GPT-5.4 Pro. You actually have to click the model selector. Don't use Auto. The prompt was almost unassuming: "Could Erdős be wrong? Could the reasoning behind this bound be flawed?" And then the model worked. Completely autonomously. 125 pages. Around 100,000 tokens. Cost: somewhere between $100 and $1,000. Reality check: tomorrow I'm flying to an oil & gas company in Hannover. Zurich → Hannover one-way: $800. So the token cost of solving an 80-year-old mathematical problem is in the order of a single business trip. 🔧 The trick: not a better screwdriver — a different wrench entirely For 40 years mathematicians attacked this with geometric tools: incidence geometry, Szemerédi-Trotter, crossing number method. Those tools hit a natural ceiling — the n^(4/3) bound. The AI did something else. It pulled a completely different key out of the toolbox: algebraic number theory. CM fields. Complex multiplication. Infinite Galois towers. It didn't solve the problem. It reformulated it — from a geometric problem to a number-theoretic one. And suddenly the answer became much more concrete. 🤖 The DeepMind counter-punch: AlphaProof Nexus + Lean Then Google DeepMind dropped the receipts. Their system AlphaProof Nexus claims to have solved: 9 open Erdős problems44 additional open conjecturesA 15-year-old problem in algebraic geometryAnd here's where it gets architectural. AlphaProof Nexus combines AI reasoning with a formal verification tool called Lean. The AI doesn't just spit out an answer — it produces a step-by-step proof, and Lean mechanically verifies every single step. Every logical leap is checked. Incorrect assumptions are rejected. The final proof meets strict mathematical standards. Cost per problem: a few hundred dollars in compute. ⚖️ Two religions: human-verified vs machine-verified This is now a genuine philosophical split in the AI math community: OpenAI's approach: let the LLM produce the proof, then send it to 9 of the world's top mathematicians — including Fields Medal winners like Noga Alon, Daniel Litt, Melanie Wood — to verify by hand. Slow. Authoritative.DeepMind's approach: let the AI prove it AND let the machine (Lean) verify it. Fast. Reproducible. But — you have to trust Lean.Both approaches address the hallucination problem: AI models can invent unproven statements, skip difficult parts, present incomplete proofs as finished. Human review and machine verification are two different solutions to the same fundamental risk. 🛑 The Hassabis caveat: AGI is still far Demis Hassabis (DeepMind CEO) reminds everyone: "For an AI, this wasn't actually that hard." The problem is extremely difficult to solve, but it's bounded. AGI would require: Creativity across multiple fields simultaneouslyIndependent reasoningOriginal idea generationToday's systems are powerful specialized tools — not minds. But here's the catch: the most clever thing the AI did wasn't the solution. It was the cross-domain reformulation. And that's exactly where your R&D department needs to wake up. 🧬 Why your R&D needs this — silos, Da Vinci, AlphaFold Pharma R&D is the textbook silo problem: Medicinal chemists define and find targetsBiologists know the pathwaysStatisticians wade through the dataThey work in their silos. They don't talk on the level where breakthroughs happen. Leonardo da Vinci could. Math + chemistry + physics + anatomy — all in one head, all connected. Today that's impossible for a human because of information overload. But an AI? An AI has exactly that cross-domain synthesis ability. Side note: Google DeepMind already won the Nobel Prize 10 years ago — for AlphaFold solving the protein-folding problem. Pure cross-domain AI. If pharma had taken that seriously, they'd be a decade ahead today. 🦴 The uncomfortable truth about your senior researchers Who are the most expensive people in any R&D department? Not the juniors. The 30-year veterans earning three-quarters of a million euros a year. And they are the worst AI users. Because they fundamentally say: "I've done research like this for 40 years. I don't need ChatGPT." When you hire a postdoc in 2026, "is he good in his domain?" is no longer the only question. The new questions: Can he prompt a reasoning model correctly?Can he ask cross-domain questions? "How would a biologist see this? How would an economist see this?"Does he click "Auto" or does he deliberately choose GPT-5.4 Reasoning?⚖️ The legal department will be your next blocker Imagine: you've found something genius with ChatGPT. You want to patent it. Who stops you first? Legal. Does it belong to us? Or to OpenAI?Does it belong to Microsoft (if you used Copilot)?Who holds the patent?The answers aren't clarified yet. Your discoveries may sit in legal review for 2 years. Plan for it. 🎯 Three Monday Actions...

    #128 - How ChatGPT Cracked an 80-Year-Old Math Problem for $1,000
  6. Jun 8

    #127 - [Quickbite] - Chief AI Academy — Sneak Peek into Session 1

    People keep asking: "What are you actually teaching in the Werchota Chief AI Academy?" So here's a Quickbite sneak peek into Cohort 1, Session 1 — the four moments that made the room go silent. Quick context: each cohort = small group (5-15 business leaders — CFOs, Chief AI Officers, heads of procurement, HR directors), 4 weekly sessions × 2 hours. Not just listening to Malcolm — also to Maria (co-founder) and Damian (associate partner, head of engineering). And critically: participants talking to each other. Because when you leave the academy, you shouldn't just talk to us — go talk to your peers. 💥 Moment 1: The Software Armageddon Opened the session with the stock charts: HubSpot down 30-40%, Gartner in free fall since Q3, Adobe Duolingo essentially dead, Salesforce bleeding. None of these companies will literally disappear tomorrow — but the rate of decline is accelerating because their customers can now build the same thing themselves. Example demoed live: a procurement participant uses IronCloud for contract reviews — €40,000/year. We showed Claude doing the same thing, more personalized, more tailored to her business, for ~$50 in tokens. The room went silent. The pushback: "But Malcolm, companies are still buying software." Yes — out of habit. As soon as they understand headless software (no UI, just an MCP server or API key, AI orchestrates it), the whole game changes. We're giving the wrong software to the wrong people right now. Stop rolling out Microsoft Copilot to your knowledge workers. Roll out Claude Code and Codex instead — with compliance-friendly options like OpenCode. Then they can build their own solutions. 🧠 Moment 2: Reverse Prompting (the killer technique) Someone said: "I tried your prompt 'make me a sexy dashboard' and nothing came out." So we went back to basics — because most people still don't know how to prompt. The technique: Reverse Prompting. Inspired by how psychologists work — they don't ask you to articulate your trauma in one sentence. They ask you questions, and through your answers, you discover what you actually want to say. Your command to Claude/ChatGPT:  "Ask me 10 questions, multiple choice, 4-5 answers each. Don't ask all at once — three at a time, wait for my answers, then formulate the next three." Now the AI is interviewing you. While you answer, you start to actually understand what you want. The AI surfaces concepts you didn't even know existed — "Do you want a Streamlit app, or maybe a web app?" — and you go: "Wait, what's a Streamlit app? Yes, that one." Promise: Malcolm will pay for your dinner if reverse prompting does not measurably improve your output. 🧠 Moment 3: The Second Brain that called me the WORST salesperson The werchota.ai "second brain" — built in 48 hours, now running in production on Azure (Victor's setup) — has access to: Every email from every employee (including mine — anyone in the company can query my inbox)Every meeting transcriptEverything on SharePointQueryable via TeamsLive demo: I asked it to do a SWOT analysis of Malcolm, focused on sales. It ran for 30 minutes. Produced a dashboard that called me out: "Malcolm is the weakest persona on the entire sales team""Rarely has an agenda going into calls""Talks all over the place, ends calls with 'OK, bye'""No follow-up discipline. No paper sent in 3 days, no follow-up in 1 week""Creates information overload for customers"The room went silent. I love when that happens. Because the point isn't to publicly roast me — it's that a Second Brain lets you do this for everyone in your team. People know their strengths. They struggle to articulate weaknesses. The Second Brain extracts them — and then Marsha can jump in on my post-meeting communication, Alex can cover ABCD, and the team plays to its actual gaps. 🦴 Moment 4: "You've been hiring AI Neanderthals" I showed them what an AI-native business leader can do — Claude Code, Codex, prompting fluency, voice-to-Excel, MCP servers, headless software integration. Then I asked: "Be honest. Your last 10 hires. How many can do this?" Answer in the room: essentially zero. You're running an AI-powered company while continuously hiring AI Neanderthals. Then you wonder why adoption is slow. If you want your company to stay Neanderthal-shaped and disappear in the next 1-2 years, continue hiring like this. I don't care. But you should. Remove "Microsoft Office" from your job descriptions. Replace with prompting, AI tool fluency, understanding of where the tech is going. The biggest leverage you have right now is who you hire next. 🎯 Three Monday Actions Try Reverse Prompting today. Take any messy goal you have, paste it to Claude/ChatGPT with the formula above. Free dinner if it doesn't work.Audit your last 10 hires against ~10 AI skills (prompting correctly, Claude Code / Codex fluency, MCP understanding, voice-to-Excel, second-brain literacy, etc). Expect 1-2 out of 10. You've been hiring problems — now they need re-training. Going forward, hire people who already use AI natively. Biggest leverage in the company.Build a Second Brain — even a small one. Don't have to go enterprise like we did. Start with a project: shared email address, project files, meeting transcripts of one initiative. Build it. Query it. Watch what surfaces.💬 What participants left with A Swiss consultant: "I realized I don't have a workload problem — I have a cross-department visibility problem. The Second Brain would solve that for me." Another participant: "I'm going to go try reverse prompting tomorrow. This alone is worth the session." That's the Chief AI Academy in 18 minutes. Want to come to Cohort 2? See the link below. ⏱️ Timestamps 00:00 — What is a Quickbite + intro to the Chief AI Academy format02:30 — Who comes: CFOs, Chief AI Officers, procurement, HR directors05:00 — Moment 1: Software Armageddon — HubSpot, Gartner, Adobe, Salesforce08:00 — IronCloud €40k/year vs Claude $50 in tokens — live demo10:00 — Headless software + MCP servers explained12:00 — Moment 2: Reverse Prompting — the dating-your-psychologist technique14:00 — Moment 3: Second Brain — the SWOT that called me the worst salesperson16:00 — Moment 4: AI Neanderthals — your last 10 hires17:30 — Three Monday actions + participant takeaways🎙️ About the Host Malcolm Werchota runs AI adoption programs for companies across Europe — close to 90 companies advised, majority in the DACH region. After 15+ years at Novartis and Schlumberger, today's focus: AI without the b******t. Lecturer at ESADE and HSLU. Studied in Leoben. 🚀 Resources for Executives 📚 Chief AI Academy — AI for Decision Makers ← Cohort 2 ...

    #127 - [Quickbite] - Chief AI Academy — Sneak Peek into Session 1
  7. Jun 5

    #126 - [AI Drama] - Your Continuity Plan Doesn't Cover Drones. It Should.

    Welcome to AI Drama. About a year ago, in a single night across five Russian airbases, 41 aircraft were destroyed. TU-95 strategic bombers. TU-22 M-3s. A-50 AWACS surveillance planes. Estimated damage: $7 billion. No air raid sirens went off. No interceptors scrambled. The attack didn't come from the sky. It came from trucks. Ordinary containers, driven inside Russian borders. The drivers were tricked — "bring the container, someone will pick it up." Nobody ever came. Inside: 117 FPV drones, loaded with explosives. The containers opened on remote command. The drones deployed one after another. Each flew only 300-400 meters to strike. Total operation cost: maybe $2-3 million. Damage inflicted: $7 billion. ROI: 3,500×. Some call it Russia's Pearl Harbor — but delivered by drones in a truck. This was Operation Spider's Web. Welcome to AI Drama. Today: the new way of running wars, teenagers in basements, and the end of "safe distance" as a concept. 🎮 Mykola, 19, gaming streamer turned drone pilot Three years ago, Mykola was a 16-year-old Counter-Strike streamer in Kharkiv with a few thousand Twitch followers. Today he's in the Ukrainian army — not because he wanted to, but because Ukraine started a new program: only soldiers between 18 and 24 can operate drones. Why 24? After 24, your brain is too slow. His "front line" is a bombed-out basement. Four laptops, three pairs of FPV goggles, a controller that looks like a PlayStation. The only light: an LED running on a battery that'll die in 2-3 hours. He flies a quadcopter that costs $300-400 in parts, 8-10 km out to a Russian position, 150 km/h. Tilts left, tilts right. Russian soldier sees it, has milliseconds to react. Impact in 4 seconds. Live feed disappears. Mykola goes to Telegram. Doesn't celebrate. Reaches for the next drone. It's only 2 PM. This is his 12th mission today. 🏭 The numbers that should terrify NATO USA: ~100,000 military drones produced per yearUkraine: 4.5 million military drones per yearRatio: 45:1 — from a country with zero drone industry four years agoIndustrial ceiling: 8-10 million drones/year projectedBloomberg: Ukraine produces more drones than the entire NATO alliance combinedUkraine has created an Unmanned Systems Forces — a military branch dedicated entirely to drone warfare. No other country in the world has this. ⚠️ Aurora 26 — when NATO learned the hard way April-May 2026, Gotland Island, Sweden. NATO exercise Aurora 26. 18,000 soldiers from 13 NATO countries. Ukraine was invited as the attacking force. The exercise had to be STOPPED THREE TIMES. Each time, the NATO troops would have been annihilated. The Ukrainian pilot said: "If it was real life, they would have all been dead." Swedish Defense Minister General Michael Claesson, after the exercise: "The fastest way for any Western force to learn about drone and counter-drone warfare is to go and listen to the Ukrainians." 🤖 Inside the $700 killing machine An FPV drone is a small quadcopter, $300-3,000. Inside lives a chip from one of our darling companies: the NVIDIA Jetson Nano (or Jetson Orin). Matchbox-sized. Costs $100-300 per unit. Available on Amazon. It wasn't built for war. It was built for robot vacuum cleaners, DIY hobbyists, AI students. NVIDIA didn't sit down and say "let's build chips for drones that kill people." It just happened. Add a Ukrainian company called Fourth Law's TFL-1 computer-vision module ($100). Now the operator only needs to fly within 400-500m of target. The Jetson takes over the final approach. Hit rate without AI: 30-50%. Hit rate with AI Jetson + computer vision: over 80%. Today, 20+ Ukrainian brigades use this AI copilot. Microsoft Copilot, but for killing people. Given to 19-year-olds. Total bill of materials: ~$700. Even priced at $10,000 (it's not), compare: $10,000 drone vs $5,000,000 Russian armored fighting vehicle$10,000 drone vs $100,000,000 Tupolev TU-95 bomber🇪🇺 Why this is YOUR problem (yes, even in DACH) Picture Werner: head of security at a mid-sized DACH automotive company. 1,500-2,000 employees. Three plants. Just-in-time delivery to Mercedes, BMW, Audi. His business continuity plan covers fire, floods, cyberattacks, supplier failure. It does not cover fiber-optic drones flying inbound from a loading dock. "But Malcolm, drones can't fly all the way to Bavaria!" You missed the point. It's $700 to build. You can build it in a garage. You can launch from a container. This was already done a year ago. And you can't send one — you send swarms. Proof? On May 30, 2026, Russia launched 800 drones in a single night. The Bundeswehr's entire drone fleet today: 500-600 units. Russia could wipe out Germany's entire drone stock in one night and still have 200 left over. 🏢 The German players quietly building this Helsing (Munich) — $12B valuation, $600M latest round led by Daniel Ek (Spotify founder). Their HX2 loitering munition uses advanced AI targeting.RF1 Resilience Factory (southern Germany) — produces 1,000+ HX2 units per month. Ukraine has ordered 10,000.Even the Bundeswehr is now buying from Helsing.🎯 Five Monday Actions for every European exec Drone-airspace continuity plan: Sit your logistics chief + ops chief + insurance + general counsel down. Ask: "What are our assumptions for business continuity if drones enter our airspace?" These plans don't exist yet.Map drone-exposed choke points in your supply chain. Belgium has had repeated airspace disruptions. Strasbourg is on the border. Don't assume "too far."Bring in someone who understands Helsing, the Quadcopters, computer vision, edge AI. You can't buy weapons for your factory — but the tech stack is migrating to commercial applications.Build AI-controllable machines. Your next product's UI shouldn't look like 1990. Build it controllable via MCP server. Currently nobody is doing this — first mover wins.Have a board conversation about it. Poland already is. US, Israel, South Korea already are. DACH boardrooms — not yet.🌐 The wild part: most of this tech is OPEN SOURCE Go to Google or Perplexity right now. Type: "GitHub repo drones". You'll find: Drone log analyzers (flight log analysis dashboards)Fully autonomous VTOL repositoriesComputer vision targeting modulesHundreds of repos with the full AI tech stackAnybody can build this right now. Not state actors. Not criminal organizations. Mykola and his mother and his sister, in a kitchen. FPV drones used to cost $50,000. Today: $500. Every new AI model + every chip generation makes them cheaper. 🎬 What this episode is really about Ukraine in the last four years has see...

    #126 - [AI Drama] - Your Continuity Plan Doesn't Cover Drones. It Should.
  8. Jun 1

    #125 - [Quickbite] - Microsoft Bans Claude Code — and Takes the Ferrari Away From Its Engineers

    Imagine your company car is a Lamborghini. Or a Ferrari — doesn't matter. You drive it to work every day. You're productive. You're happy. And then your CEO walks in and says: "Starting next month, you're driving a Skoda Octavia." That's exactly what just happened at Microsoft. And it affects you directly — even if you've never written a line of code. Last week, May 14, 2026, an internal memo landed at Microsoft's Experiences and Devices Division. Windows, Microsoft 365, Outlook, Teams. Tens of thousands of engineers. The memo came from Rajesh Jha, Executive Vice President. The content in one sentence: We're shutting down Claude Code. Deadline: June 30, 2026. The absurdity: six months ago — December 2025 — Microsoft aggressively rolled out Claude Code to those same engineers. Thousands of seats. Even designers and project managers got access. The original ask: install this, experiment, build prototypes. Why the reversal? Not because Claude Code is bad. Because it's too good. It was better than Microsoft's own tool — GitHub Copilot — at exactly the work that matters: multi-file refactoring, architectural work, rapid prototyping. Microsoft sells GitHub Copilot to the world as its AI developer flagship. Microsoft invested $13 billion in OpenAI. And for six months, Microsoft's own engineers quietly preferred a competitor's product from Anthropic. That's not embarrassing — that's a strategic bomb. 📊 What separates Claude Code from GitHub Copilot Copilot is autocomplete. You type, Copilot suggests the next line. You're driving. Passive. Like a Skoda with cruise control.Claude Code is agentic coding. You say: "Build me an app that recognizes my Sonos speakers and starts music when my Tesla arrives home." Claude works two, three, even seven hours autonomously. Reads the whole codebase. Refactors. Tests its own output. You're no longer driving — you're a project manager.Context window: 1 million tokens (rumored 12M coming). The AI's brain fits the entire codebase.Extended thinking: Claude stops, plans, reasons, will tell you when something is nonsense. Copilot codes blindly forward.Multi-file autonomy: Claude grabs "helper" agents and works in parallel across the codebase.💸 The pricing question Claude Code Enterprise: $150 per seat per month. GitHub Copilot: $10 to $30. Microsoft engineers were using the 10× more expensive tool — and when they ran out of tokens, they paid out of their own pocket for more. Like a free-to-play game, except here the tokens produce production code. ⚠️ The Amazon precedent Microsoft is not the first to make this mistake. End of 2025 Amazon banned Claude Code and Codex internally and mandated their in-house tool "Kiro." What happened immediately? A 13-hour AWS outage in China. Engineers stuck with a Skoda Octavia facing a Ferrari-sized problem. By April 2026, Amazon reversed course and re-enabled Claude Code. Google does something similar: Claude Code is blocked by default — except at DeepMind, their top AI division. SpaceX just paid $60 billion for an option on Cursor (a Claude Code competitor). The pattern is identical everywhere. 🇪🇺 The DACH / European lesson If you're a CTO, VP of Engineering, or founder in a typical European tech company: your developers are already using these tools. As shadow AI. On personal subscriptions. Quietly in the evenings. Here's how to figure that out — without any survey: Two years ago: ~3,000 lines of code per developer per dayWith Copilot: jump to 6,000–9,000 (2–3×)With Claude Code: jump to 30,000–300,000 (10–100×)Just look at the output. That's your audit. Done in a Monday morning. 🇪🇺 The sovereign alternative If data sovereignty matters: Mistral Codestral — 22B-parameter code model, 80+ programming languages, EU infrastructure, GDPR-native. Mistral just raised nearly $1 billion from European banks to build exactly this. Plus the upcoming Cohere-Aleph Alpha merger (Schwarz Group, €500M) explicitly building for DACH enterprises. You don't have an excuse anymore. 🏭 The hackathon moment Three days ago we co-ran a hackathon at a major German manufacturing company. 20 top developers in the room with the absolute best tools — OpenCode, Open Terminal, Claude Code. Phenomenal. But then the question: 20 people at the table, 6,000 in the corporation. When do the other 5,980 get the same tools? 🚀 How we work at werchota.ai Every single person at our company uses Claude Code. 85% of all our work is done by Claude Code and AI agents. Porni (journalist) — Claude Code. Alex (finance) — Claude Code. Not because they code. Because the tool has become universal. 📌 Three Monday actions Shadow AI audit. Look at code output per developer across 2 years. Who made the 10× jump? That person is secretly using Claude or Codex.A/B test with a real task. Same task, same 24 hours. One developer "old way," one with Claude Code. Compare output, error rate, completeness.Three-tier data classification. Tier 1 non-sensitive = any tool. Tier 2 internal business logic = EU-hosted (Mistral). Tier 3 regulated data = security review. Not a ban. A policy.🎬 The bigger question Microsoft will reverse this in 2-3 months. Just like Amazon did. But you have a more important problem: are you keeping the Ferrari away from your engineers — or finally giving it to everyone? ⏱️ Timestamps 00:00 — Cold open: The Lamborghini, the Microsoft memo, the June 30 deadline03:00 — Agentic coding vs. autocomplete — the two worlds05:30 — Context window, extended thinking, multi-file autonomy07:00 — The $150-vs-$20 question and why engineers still pay09:00 — Amazon's 13-hour AWS China outage + Google + SpaceX-Cursor11:00 — How to audit your shadow AI in 5 minutes13:00 — Mistral Codestral + Cohere-Aleph Alpha as the sovereign alternative14:30 — The hackathon: 20 vs. 6,000 — the question every CTO must answer15:30 — werchota.ai: 85% Claude Code, every single person16:00 — Three Monday actions + close from Bregenz🎙️ About the Host Malcolm Werchota runs AI adoption programs for companies across Europe. After 15+ years at Novartis and Schlumberger, today's focus: AI without the b******t. Last week live at the AIM Summit in London — after Lord Melvin (former Chief of the Bank of England) and before Eric Trump, in front of 150 investors. Lecturer at ESADE and HSLU. Studied in Leoben. 🚀 Resources for Executives 📚 Chief AI Academy — AI for Decision Makers👥 AI...

About

Malcolm Werchota's AI Cookbook Show is where artificial intelligence meets authentic business transformation. Known for his direct style and willingness to show AI in action—even during live presentations—Malcolm helps organizations understand that AI isn't about replacing humans but amplifying their capabilities. From voice-note productivity hacks to real-time meeting intelligence, this podcast delivers actionable insights for immediate implementation.