The Generative AI Meetup Podcast

Mark and Shashank

Hosted by Mark and Shashank, software engineers and organizers in Silicon Valley. Get their grounded perspective each week as they explore the generative AI landscape through news analysis, tech discussions, hands-on experiments, and clear explanations. Dive into the latest language models, AI agent capabilities, and RAG techniques. Understand the hardware race, key research, startup trends, benchmarks, and the real-world impact of AI across industries like healthcare, robotics, and creative work. We also test AI limits, explain core concepts, discuss ethics, and interview builders shaping the field. For engineers, developers, researchers, and anyone seeking a practical understanding of AI’s rapid evolution and its applications.

  1. Jun 23

    What happened to my Fable?

    https://novacut.ai/  Description: Anthropic pulls access to Fable, and China responds the same day with GLM 5.2. In this episode we break down the escalating AI arms race, US export controls on chips and frontier models, and whether the "Great Firewall of America" is already here. ⏱️ Topics: Anthropic restricts Fable — what happened and why China's GLM 5.2 release and how close they're catching up US trust, surveillance, and AI gatekeeping Token pricing chaos — cost per task vs. cost per token Model routing, loop engineering, and autonomous agents Anthropic's Mythos model and Fable safeguard philosophy Xiaomi NEMO V2.5 Pro Ultra Speed Midjourney's bizarre health spa pivot AI Engineer Conference wrap-up 🔗 Links & Resources:  Fable   https://www.anthropic.com/news/claude-fable-5-mythos-5  https://www.theregister.com/security/2026/06/15/feds-freaked-over-fable-5-after-simple-fix-this-code-prompt-not-jailbreak-says-researcher/5255827 “Fix this code” https://support.claude.com/en/articles/14328960-identity-verification-on-claude    Midjourney https://www.midjourney.com/medical/blogpost  Full body ultrasound CT scanner   Xiaomi 1000tps https://mimo.xiaomi.com/blog/mimo-tilert-1000tps (MiMo-V2.5-Pro-UltraSpeed: Pushing 1T-Parameter Model Generation Speed to 1000 TPS Best opensource model https://z.ai/blog/glm-5.2 📌 Timestamps in the chapters section above. #AIPodcast #Anthropic #Fable #GLM52 #AIArmsRace #LLM #GenAI 0:00 Intro: Anthropic restricts Fable access 1:00 China's response and GLM 5.2 release 1:55 US trust and AI model reliability 2:37 Geopolitics and AI regulations 4:17 AI arms race and export control limits 5:49 Fable usage experience and value 6:30 Z.AI subscription and pricing comparison 9:07 Subscription limits vs. API usage 10:30 GLM 5.2 token limits and utility 13:52 China catches up: GLM vs. US models 15:06 Intelligence index and model cost trend 17:40 Token pricing complexity and value 23:30 Cost per task vs. cost per token 24:35 Model routing and usage optimization 29:43 Loop engineering and autonomous agents 33:23 NovaCut: AI video editor and ad loops 37:07 Anthropic's Fable re-release timeline 38:41 US gatekeeping and China's advantage 40:11 Great Firewall of America risks 44:07 Mass surveillance and free speech 46:01 Global AI trust and market shift 47:48 Enforcing identity checks on AI 48:54 Export bans on chips and hardware 51:52 US restricts allies from frontier models 52:39 Anthropic's talent and Mythos model 53:51 Fable safeguards and US government view 1:00:12 Xiaomi NEMO V2.5 Pro Ultra Speed model 1:02:39 Optimizing for intelligence, cost, speed, and size 1:13:19 Midjourney's unexpected health spa pivot 1:28:30 AI Engineer conference and podcast wrap

    1h 30m
  2. Jun 7

    The Best Open Source US Model (Right behind China)

    https://novacut.ai/  https://genaimeetup.com/  Anthropic has officially closed a $65 billion Series H at a $965 billion valuation, nearly 2.5x its valuation from just 100 days ago. Meanwhile, funding is flowing across the ecosystem: Frameworks AI at $15B, Baseten at $11B, OpenRouter's $113M Series B, and Cognition AI's $1B Series D. NVIDIA went on an open-source super week with Nemotron 3 Ultra, Cosmos 3, and Nemotron 3.5 ASR. Microsoft dropped 5 new MAI models. Google released Gemma 4 12B, and Anthropic shipped Opus 4.8. On the benchmarks front, DeepSWE crowns GPT-5.5 as the leader in long-horizon coding tasks, while ITBench shows even frontier models struggle with real-world SRE incidents — Claude Opus 4.7 tops out at just 47%. Plus: Cloudflare acquires VoidZero to build the future of AI-native edge development, and Google is paying SpaceX $920M/month for compute. Topics covered: • Anthropic's $65B Series H and path to $1T • Fireworks AI, Baseten, OpenRouter & Cognition funding rounds • Microsoft's 5 new MAI models • NVIDIA's open-source super week (Nemotron, Cosmos 3) • MiniMax M3, Gemma 4 12B, JetBrains Mellum2, Opus 4.8 • DeepSWE benchmark: GPT-5.5 leads long-horizon coding • ITBench: Frontier models under 50% on real SRE tasks • Cloudflare + VoidZero for AI-native edge dev • Google's $920M/month SpaceX compute deal #AI #Anthropic #NVIDIA #OpenAI #AInews #TechNews #LLM     Funding rounds Anthropic formally confirmed the closure of its $65 billion Series H funding round at a post-money valuation of $965 billion. This represents a 2.5-fold increase over its $380 billion Series G valuation from February 2026, adding $585 billion in value in approximately 100 days https://www.anthropic.com/news/series-h  Frameworks AI raising at 15B valuation representing a near fourfold increase from its $4 billion Series C valuation recorded in October 2025 processing 15 trillion tokens daily for major production clients including Cursor, Notion, and Perplexity https://finance.yahoo.com/sectors/technology/articles/fireworks-ai-eyes-15-billion-174609357.html Baseten is raising 1B at 11B valuation annualized revenue, which skyrocketed from $200 million to $600 million over a single quarter https://techstartups.com/2026/05/26/ai-inference-startup-baseten-in-talks-to-raise-1-billion-at-11-billion-valuation/  OpenRouter has secured a $113 million Series B funding OpenRouter has experienced exponential traffic growth, with weekly production throughput expanding fivefold from 5 trillion to 25 trillion tokens over a six-month horizon https://www.businesswire.com/news/home/20260526953416/en/OpenRouter-Raises-%24113-Million-CapitalG-led-Series-B-as-Weekly-Volume-Explodes-to-25T-Tokens  Further up the stack: Cognition AI secured a $1 billion Series D round led by Lux Capital and 8VC https://cognition.ai/blog/series-d   Model Releases MAI models: MAI-Code-1-Flash: A 5-billion active parameter model optimized for ultra-low latency within GitHub Copilot and VS Code. MAI-Image-2.5: A high-fidelity image generation model ranking third on global image evaluation arenas, outperforming competing architectures like Nano Banana Pro. MAI-Transcribe-1.5: A multi-lingual speech processing engine offering fivefold speed improvements across 43 languages. MAI-Voice-2: Natural audio and voice generation across 15 languages, available at a highly competitive price point. Web IQ: A search-grounding API engineered to directly compete with Perplexity. https://microsoft.ai/models/    https://www.peoplematters.in/news/ai-and-emerging-tech/uber-imposes-dollar1500-monthly-ai-spending-limit-on-employees-amid-rising-costs-50073    Nvidia has executed an "Open-Source Super Week," positioning itself as a dominant software and model publisher: Nemotron 3 Ultra (best US open source open weights model but behind china): A massive 550-billion parameter MoE (55 billion active) designed with a 1-million token context window, optimized specifically for high-throughput, cyclical agent loops. It achieved peak throughput rates of 400 tokens per second on day-zero optimized clusters. Cosmos 3: A physical AI world-modeling framework comprising 16-billion Nano and 64-billion Super variants. Built on a Mixture-of-Transformers (MoT) architecture, Cosmos 3 natively binds textual, visual, auditory, and physical kinetic vectors. Nemotron 3.5 ASR: A highly compact 0.6-billion parameter streaming speech recognition model pushing sub-100 millisecond latencies across 40 language locales.   https://www.minimax.io/models/text/m3  MiniMax M3: A 1-million token context model hitting 59.0% on SWE-Bench Pro and 74.2% on MCP Atlas, though noted for high token consumption due to intensive internal self-validation loops.   https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12b/  Gemma 4 12B: Google's Apache 2.0 on-device model, which utilizes an encoder-free architecture that projects vision and audio vectors directly into the text-token space, bypassing separate CLIP-style encoders to minimize local memory footprints. https://www.jetbrains.com/mellum/  JetBrains Mellum2: A compact 12-billion parameter MoE (2.5 billion active) engineered for ultra-low latency routing and retrieval-augmented generation (RAG) sub-agents within developer IDEs. Opus 4.8 https://www.anthropic.com/news/claude-opus-4-8    https://www.cnbc.com/2026/06/05/google-to-pay-spacex-920-million-a-month-for-xai-compute-capacity.html      Benchmarks: https://deepswe.d atacurve.ai/blog https://venturebeat.com/technology/deepswe-blows-up-the-ai-coding-leaderboard-crowns-gpt-5-5-and-finds-claude-opus-exploiting-a-benchmark-loophole (GPT 5.5 the winner in long horizon tasks) a highly complex software engineering benchmark focused on original, long-horizon tasks across five distinct programming languages. Comprising 113 chaotic tasks across 91 live, production-grade repositories, DeepSWE forces agents to generate 5.5 times more code and modify an average of 7 separate files per task compared to standard evaluations. On this challenging leaderboard, GPT-5.5 leads with a score of 70%, establishing a significant 16-percentage-point lead over contemporary alternatives I think older benchmarks where models reach ~90% accuracy can be considered saturated. Few percentage points don’t give us any good signal.  https://research.ibm.com/publications/developing-ai-agents-for-it-automation-tasks-with-itbench  ITBench-AA, an evaluation framework focusing on live Kubernetes incident response and Site Reliability Engineering (SRE) operations. Comprising 59 live, containerized SRE incident snapshots, the results are remarkably sobering: every frontier model scored under 50% on successful incident resolution, with Claude Opus 4.7 leading at 47% and GPT-5.5 following closely at 46%.   Edge AI announcements: https://www.cloudflare.com/press/press-releases/2026/cloudflare-acquires-voidzero-to-build-the-future-of-the-ai-native-web/  The consolidation of the AI-native developer stack has reached the runtime virtualization layer. Cloudflare recently completed the acquisition of VoidZero, the development group responsible for Vite, Vitest, Rolldown, and Oxc, backing the transaction with a $1 million open-source ecosystem fund. This acquisition is highly strategic; as autonomous agents write an increasing proportion of production software, local development environments, compilation pipelines, and bundlers must be optimized for execution speeds that match agent speeds. Cloudflare's goal is to construct a localized, full-stack edge playground. In this sandbox, AI agents can generate, test, bundle (utilizing the highly parallelized, Rust-based Oxc and Rolldown engines), and deploy entire web applications end-to-end within milliseconds. This architecture completely bypasses traditional local machine container bottlenecks, enabling high-velocity agent loops to execute in a fully sandboxed, web-scale edge runtime.

    1h 55m
  3. Mar 5

    AI Matches Human Intelligence, Pentagon Drama, and the Rise of Agent Swarms

    Youtube Channel: https://www.youtube.com/@GenerativeAIMeetup Mark's Travel Vlog: https://www.youtube.com/@kumajourney11 Mark's Personal Youtube Channel: https://www.youtube.com/@markkuczmarski896 Attend a live event: https://genaimeetup.com/ Shashank Linked In: https://www.linkedin.com/in/shashu10/  Novacut: https://novacut.ai    Mark and Shashank break down the latest developments in AI from their travels in Fukuoka and Seychelles. They cover Gemini 3.1 Pro matching human performance on the ARC-AGI-1 benchmark at a fraction of the cost, the upcoming ARC-AGI-3 video game-style test, and why only three US companies (OpenAI, Anthropic, Google) seem to be pushing state-of-the-art right now while Meta and xAI deal with leadership shakeups. The conversation moves to OpenAI's GPT 5.3 Codex Spark model running on Cerebras hardware for lightning-fast inference, Abu Dhabi's M42 initiative sequencing 700,000+ genomes and centralizing health records for AI-driven healthcare, and the viral OpenClaw incident where an AI agent wrote a hit piece on a human open-source maintainer who rejected its pull request. They also discuss the Anthropic vs. Pentagon drama over autonomous weapons and mass surveillance restrictions, an ex-Google Maps PM who vibe-coded a Palantir-style intelligence dashboard in a weekend, and their hands-on experiences with Claude Code, Codex, Cursor, and MCP integrations. The episode wraps with thoughts on agent swarms, the human-in-the-loop problem for taste-driven tasks, and whether we're close to the first solo-founder billion-dollar company powered entirely by AI agents.

    1h 39m
  4. Jan 6

    Groq, Hotel Delivery Robots, and Mark Launches a Company

    It’s been a travel-heavy hiatus—Mark’s been living in Spain and Shashank’s been bouncing across Asia (including a month in China)—but they’re back to unpack a packed week of AI news. They start with the headline hardware story: the Groq (GROQ) deal/partnership dynamics and why ultra-fast inference is becoming the next battleground, plus how this could reshape access to cutting-edge serving across the ecosystem. From there, they pivot to NVIDIA’s CES announcements and what “Vera Rubin” implies for data center upgrades, cost-per-token curves, and the messy real-world math of rolling hardware generations. Shashank then brings the future to life with on-the-ground stories from China: a Huawei “everything store” that feels like an Apple Store meets a luxury dealership, folding devices that look straight out of sci-fi, and a parade of robots—from coffee bots to delivery robots that can ride elevators and deliver to your hotel room. They also touch on companion-style consumer robots and why “cute” might be a serious product strategy. Finally, Mark announces the launch of Novacut, a long-form AI video editor built to turn hours of travel footage into a coherent vlog draft—plus export workflows for Premiere, DaVinci Resolve, and Final Cut. They close by talking about the 2026 shift from single model calls to “agentic” systems, including a fun (and slightly alarming) lesson from LLM outcome bias using poker hand reviews. Topics include: Groq inference, NVIDIA + CES, Vera Rubin GPUs, GPU depreciation math, China robotics, Huawei ecosystem, hotel delivery bots, companion robots, Novacut launch, Cursor vs agent workflows, and why agents still struggle with sparse feedback loops. Link mentioned: Novacut — https://novacut.ai

    1h 2m
5
out of 5
11 Ratings

About

Hosted by Mark and Shashank, software engineers and organizers in Silicon Valley. Get their grounded perspective each week as they explore the generative AI landscape through news analysis, tech discussions, hands-on experiments, and clear explanations. Dive into the latest language models, AI agent capabilities, and RAG techniques. Understand the hardware race, key research, startup trends, benchmarks, and the real-world impact of AI across industries like healthcare, robotics, and creative work. We also test AI limits, explain core concepts, discuss ethics, and interview builders shaping the field. For engineers, developers, researchers, and anyone seeking a practical understanding of AI’s rapid evolution and its applications.

You Might Also Like