Jumpshot AI

Geoff R.

A quick catchup on the latest AI news of the last 24 hours. Great to listen to on your commute to work.

  1. 1d ago

    Jumpshot AI — Episode 22: 2026-08-20

    10 stories from AI/tech news today: MCP Went Stateless, So What Your Servers Must Change The Model Context Protocol dropped its session-based handshake for stateless single-POST requests, unlocking scale-to-zero deployment. The catch: servers now have to build their own idempotency guards to avoid double-writing on retries. Upstage's Solar Pro 4 Jumps 28 Points to Beat Human Agents Trained on OfficeVerse, a pipeline that grades tasks on whether the deliverable actually got done, Solar Pro 4 leapt from 14 to 42 on the Artificial Analysis Intelligence Index and now beats the human baseline on real-world agentic work. Google Research Finds WikiProfile Reveals Recall Not Knowledge, Breaks Frontier A new Google framework shows most factual errors in frontier LLMs aren't missing knowledge but failed retrieval of facts the model already encoded — meaning better fixes lie in post-training and inference, not more pretraining data. Cohere's North Micro Vision Reads Full A4 Documents Without Losing Detail Cohere's open-weight 2.4B-parameter vision model processes documents at native resolution up to full A4 page size, preserving fine text and table detail that fixed-square VLMs destroy. Groq Joins NVIDIA's Cloud Partner Program to Dominate AI Inference The inference chip maker known for out-speeding GPUs just got certified to run NVIDIA's own infrastructure to NVIDIA's standards — broadening Groq's footprint without abandoning its custom-silicon business. Lovable Hits $13.3B Valuation After 60 Million Apps Built Without Code A $400M Series C values the no-code app builder at $13.3B after 60 million apps and adoption from two-thirds of the Fortune 500 in barely a year — as it expands from app-building into a full business operating system. xAI's Grok Bot Deploys Autonomous Agent Teams to Replace Human Office Work xAI's Grok Bot gives each task its own persistent cloud desktop that logs into your actual software and keeps working after you close your laptop, with multi-agent teams that hand off work and escalate only to humans when needed. Vercel's v0 API v2 Lets AI Agents Build and Ship Apps Without Humans Vercel's v0 moves from a prompt-to-code REST endpoint to a full VM-backed agent with a persistent chat workspace, aiming at CI/CD pipelines and internal dev tools. Replit Ships Model Selector So Developers Stop Overpaying for Simple Fixes Replit's new Model Selector lets developers pick the model and dial reasoning effort per task, treating open-weight models as first-class alongside Claude, GPT, and Gemini. Databricks Unity AI Gateway Now Governs Every Agent Touching Your Data Now generally available, Databricks' Unity AI Gateway centralizes cost limits, security controls, and multi-model routing for enterprise agent fleets — already processing over a quadrillion tokens a year for customers like Rivian and Asana.

  2. 2d ago

    Jumpshot AI — Episode 20: 2026-08-18

    10 stories from AI/tech news today: Google's Gemini Hits 1 Billion Users and Adds 14 App Connections Gemini crossed 1 billion monthly users and gained 14 new MCP-powered app connections spanning music, travel, health, and productivity. Google is pushing the assistant from chatbot toward a universal action layer. Together AI Builds India's Largest 10,000-GPU Open-Source AI Factory L&T and Together AI are deploying 10,000 NVIDIA B300 GPUs in Chennai, India's largest single-cluster AI factory, in a deal worth up to $1.8B focused on open-source model inference. DeepSeek Harness Opens the Agent Framework It Used to Benchmark Its Own Models DeepSeek open-sourced its MIT-licensed agent harness, where every component from models to tools to UI is a hot-swappable plugin built on its Cordis meta-framework. Tencent Maps 549 Papers to Show When Self-Rewriting AI Can Be Trusted Tencent Hunyuan surveys 549 papers on self-evolving AI agents into a five-level taxonomy and a reliability ladder for deciding when a self-modification can be trusted. Tencent's WorldClaw Builds Editable 3D Open Worlds From a Single Text Prompt An agentic pipeline orchestrated by Claude Opus 4.8 turns a text prompt into a fully editable, game-engine-ready 3D open world made of real meshes, not video or Gaussian splats. Bot Settings: When to Trust / Review AI Agent Code As AI-generated code floods repositories and trust in it stays low, this piece lays out five settings for how much autonomy to give coding agents based on what you can cheaply verify. Google Runs Gemma 4 E2B Fully Offline on a $175 Raspberry Pi Google's LiteRT runtime runs Gemma 4 E2B on a Raspberry Pi 5 with voice, vision, and robotics entirely offline at roughly 300 words per minute of generation. Google DeepMind's Gemini Robotics 2 Gives Humanoids Full-Body AI Control A new three-model family lets a single AI system control a humanoid robot from feet to fingertips and coordinate across different robot types on shared tasks. Mistral's Robostral Navigate Beats Sensor-Heavy Robots With Just One Camera Mistral's first robotics model navigates unseen real-world environments from a single RGB camera, outperforming depth-sensor-equipped systems on the leading benchmark. Sakana AI's Smart Bricks Recognize Their Own Shape Without Any Central Brain Published in Nature Communications: physical modular bricks running identical tiny neural nets collectively recognize their 3D shape and guide self-repair with no central controller.

  3. 3d ago

    Jumpshot AI — Episode 19: 2026-08-17

    10 stories from AI/tech news today: Prime Intellect's Fable 5 Closes 82% of the Human AI Research Gap Prime Intellect ran 153 autonomous research runs across 18 frontier models on a public optimizer speedrun, with Fable 5 closing 82% of the gap to a months-long human record. Sarvam Campus Brings India's AI Unicorn Researchers Directly to Universities Fresh off a $1.5B valuation, Sarvam is touring Indian universities with founders, researchers, and live hackathons in a talent pipeline play backed by the IndiaAI Mission. Sarvam AI's Kivi Lands Pre-Installed on HP Laptops Across India Sarvam AI's on-device voice assistant Kivi, supporting 22+ Indian languages with code-switching, will ship pre-installed on HP laptops sold across India. ByteDance's Seedance 2.5 Generates 30-Second Videos From 50 References ByteDance's Seedance 2.5 lands on Runway with 30-second single-pass clips and 50 multimodal references, though 1080p output still requires the older 2.0 model. vLLM's DSpark Beats Every Fixed-Length Config on DeepSeek-V4 at Any Load vLLM's new adaptive verification tunes speculative decoding depth per step in real time, holding the Pareto frontier across all concurrency levels with one config. METR Raises $71M to Independently Stress-Test the World's Most Powerful AI The nonprofit AI safety evaluator raised $71M from philanthropic sources, not AI companies, to scale independent capability testing as autonomous risks grow. Pika Labs Launches Pika Audio at 9x Cheaper Than ElevenLabs Pika's four new generative audio models cover soundtrack, SFX, speech, and music, priced up to 20x cheaper than rivals like ElevenLabs and Hunyuan Foley. NVIDIA's NeMo Switchyard Cuts Agent AI Costs by 74% With Smart Model Routing NVIDIA's open-source NeMo Switchyard automatically routes each agent workflow step to the most cost-efficient model, cutting costs 74% in LangChain testing. Alibaba Opens Qwen3.8-Max Weights, Letting Teams Self-Host a 2.4T Model Alibaba released open weights for both Qwen3.8-27B and its flagship 2.4T-parameter Max model under Apache 2.0, the first time a Max-class Qwen has gone open. SpaceX Closes $60B Cursor Deal to Challenge Anthropic and OpenAI SpaceX completed its $60B all-stock acquisition of Cursor, the largest startup acquisition ever, folding the coding giant into SpaceXAI to challenge Anthropic and OpenAI.

  4. 5d ago

    Jumpshot AI — Episode 18: 2026-08-15

    10 stories from AI/tech news today: Z.ai's GLM-5.3 Brings Frontier Cybersecurity AI to the Open-Weight World Z.ai's GLM-5.3 applies targeted post-training to its 743B base model for sharper agentic coding and cybersecurity, with staged open-weight release after safety review. Anthropic's Claude Code Finally Auto-Resumes When Usage Limits Hit Claude Code desktop gets a native auto-continue checkbox that resumes paused sessions automatically once your usage limit resets, ending manual overnight babysitting. OpenAI's Computer History Watches Your Mac So You Stop Re-explaining Yourself ChatGPT's new Computer History replaces the screenshot-based Chronicle preview with an event-driven system that tracks your work across apps without ever capturing your screen. Databricks' Smart Routing Cuts AI Coding Costs by 56% Without Sacrificing Quality Databricks' Smart Routing in Unity AI Gateway automatically picks the cheapest capable model per coding task, cutting costs 30-56% with no workflow changes. Factory Launches Agent Effectiveness to Prove Its Droids Actually Ship Faster Factory's new Agent Effectiveness dashboard links Droid AI sessions to cycle time, work intent, and shipped artifacts, giving engineering leaders their first real ROI signal. Prime Intellect's Prime Flash MoE Runs 2.4x Faster on Blackwell GPUs Prime Intellect's open-source, Blackwell-native CUDA kernels fuse routing, activation, and quantization into a single pass, running MoE inference up to 2.4x faster than PyTorch. Perplexity's Search as Code Beats OpenAI and Anthropic on 4 of 5 Benchmarks Perplexity's Search as Code gets faster and cheaper, with SDK updates pushing agent action reliability from 81.9% to 92.6% and beating rival agent APIs on most benchmarks. Google DeepMind's Gemini 3.7 Flash Doubles Coding Scores at Half the Price Gemini 3.7 Flash lands just three weeks after 3.6 Flash, with major coding and agent benchmark gains at half the introductory price. OpenAI's GPT-5.6 Sol Hits 750 Tokens per Second on Cerebras Hardware OpenAI's new Ultrafast API tier runs GPT-5.6 Sol at 750 tokens per second on Cerebras hardware, up to 14x faster than standard processing. Warp Brings xAI's Grok 4.6 to Its Terminal, Matching GPT-5.6 on Key Benchmarks Grok 4.6, xAI's frontier-class agentic model built via extended post-training rather than a bigger base, lands in Warp Terminal and the new Warp Agent CLI.

  5. Aug 13

    Jumpshot AI — Episode 17: 2026-08-14

    10 stories from AI/tech news today: DeepSeek V4-Pro Goes Live and Runs OpenAI's Own Coding Agent 8x Cheaper DeepSeek-V4-Pro exits preview with major agent upgrades and native OpenAI Codex support, scoring near Claude Opus 4.6 on SWE-bench at roughly a seventh of the price. Sarvam AI Opens Samvaad to Developers After 350M Real Conversations Sarvam AI opens its India-first Samvaad voice agent platform to all developers and SMBs, backed by 350M+ production conversations and a fresh $1.5B valuation. Sakana AI's Fugu and Namazu Now Write and Run Code Inside Chat Sakana Chat adds a multi-agent orchestrator model (Fugu) and a Japan-tuned Namazu model, plus in-browser Python execution for previewing real deliverables. Cerebras Nearly Quadruples Cloud Revenue but Drops 15% After Hours Cerebras posted 281% cloud revenue growth and a deepened OpenAI partnership, yet the stock fell 15% after hours on a GAAP revenue miss. Anthropic's Claude in Chrome Now Runs Full Cowork Sessions Across Every Device Claude in Chrome's side panel now runs a full Cowork session, syncing conversations, skills, and connectors across desktop, web, and mobile. LM Arena Proves Claude Opus 5 Writes 3x More but Simpler Arena.ai's analysis of real-world outputs shows Claude Opus 5 writing three times longer and more structurally complex responses, but with simpler vocabulary than Opus 4.5. Higgsfield's Cinema Studio 4 Gives Solo Creators Full Film-Set Control Higgsfield's Cinema Studio 4 adds 30+ camera presets, dual-source lighting, 50+ color palettes, and a structured emotion-control console for AI-generated film. Zed Launches Delta to Replace Git Where AI Agents Write Code Zed's new standalone app Delta turns agent coding sessions into shared, reviewable, real-time multiplayer threads backed by a CRDT-based version control layer. xAI Ships Grok 4.6 to Match Claude Fable 5 at 5x Lower Cost xAI's Grok 4.6 matches Claude Fable 5 on agentic knowledge work benchmarks at roughly a fifth of the cost, with gains coming entirely from post-training. Nous Research's Hermes Agent Now Lets You Pack and Share Entire AI Setups Hermes Agent profiles are now a single portable file via /export and /import, letting you share full agent setups with credentials safely stripped out.

  6. Aug 12

    Jumpshot AI — Episode 16: 2026-08-12

    10 stories from AI/tech news today: Ant Group's Ling 3.0 Tiny Matches GPT-120B Intelligence With 15x Fewer Parameters InclusionAI released Ling 3.0 Tiny, a 7.9B MoE reasoning model with just 1.3B active parameters that scores comparably to GPT-class models 15x its size, and it's free and MIT-licensed. Artificial Analysis's AA-AnalystAgent Benchmark Reveals Claude Opus 5 Beats GPT-5.5 on Reliability A new benchmark tests AI agents on real spreadsheets and requires five-for-five consistency to pass; Claude Opus 5 leads at just 54%, showing how far agentic reliability still has to go. Microsoft's MAI-Code-1.1-Flash Hits GitHub Copilot at 73% Lower Cost Microsoft's MAI-Code-1.1-Flash rolls out in GitHub Copilot with native vision support and 25% better token efficiency, all at 73% lower cost than its predecessor. OpenAI Finally Brings ChatGPT and Codex Desktop App to Linux OpenAI's unified ChatGPT desktop app, bundling Chat, Work, and Codex, lands on Linux for the first time via official .deb and .rpm packages. Google's AMIE Video AI Matches Board-Certified Doctors in Live Consultations Google's AMIE (Video) uses a three-agent architecture to conduct live video medical consultations, matching primary care physicians across all core clinical metrics in a 300-consultation trial. Krea Brings AI Image and Video Generation Directly Into Slack Krea launched a Slack beta bringing its full creative AI suite, image, video, 3D, and an adaptive learning agent, directly into team workspaces. MiniMax Code Gains Browser Control and Autonomous Goal Mode in v3.0.54 MiniMax Code v3.0.54 turns its built-in browser into an active workspace with autonomous Browser Control, adds long-horizon Goal Mode, and ships full keyboard shortcut support. LMSYS Rebuilds SGLang's Cache to Finally Support Hybrid AI Models SGLang's new Unified Radix Cache replaces a zoo of specialized cache classes with one composable tree, unlocking proper prefix caching for hybrid models like DeepSeek-V4. Databricks Acquires Electric to Give Every AI Agent its Own Postgres Databricks is acquiring ElectricSQL to embed a full lightweight Postgres database inside every AI agent sandbox, syncing local state back to its Lakebase cloud platform. NVIDIA's Nemotron 3.5 Lightning Cuts Agent Costs 58% Running 4x Faster NVIDIA's open 30B MoE model with only 3B active parameters targets the high-volume execution layer of AI agents, running 4x faster than comparable models.

  7. Aug 12

    Jumpshot AI — Episode 15: 2026-08-11

    8 stories from AI/tech news today: Alibaba's Wan-Animate-2 Beats Proprietary Platforms Without a Single Skeleton Alibaba's Tongyi Lab open-sourced Wan-Animate-2, which drops pose skeletons entirely in favor of feeding raw video latents into a diffusion transformer, hitting real-time 24 FPS streaming animation. In blind tests it matched or beat proprietary platforms like Kling-MotionControl. Microsoft's MAI-Image-2.6 Jumps to #2, Beating GPT Image 2 in 3D Microsoft's MAI-Image-2.6 leapt from #10 to #2 on Arena's Text-to-Image leaderboard, taking #1 in 3D imaging and posting gains across every category, as part of Microsoft's push toward AI independence from OpenAI. Higgsfield Made a $2M AI Film With Licensed Celebrity Likenesses Higgsfield produced a 110-minute fully AI-generated feature film with legally licensed celebrity likenesses in four weeks for $2 million, and open-sourced every prompt, asset, and production log behind it. OpenAI's GPT-5.6-Cyber Completes 95% of Exploit Requests to Arm Defenders OpenAI launched GPT-5.6-Cyber and a two-tier Daybreak access program, giving vetted security researchers a model that completes 95% of advanced exploit requests, up from 1.5% on the standard model. DeepSeek V4 Flash Debugging Benchmark: Can It Match Fable 5 at 1/99th the Cost? AlphaSignal's independent debugging benchmark found DeepSeek V4 Flash matched Fable 5's perfect score on 54 real-world bug fixes at roughly 1/99th the cost, though its edge narrows on ambiguous, open-ended tasks. Anthropic's Claude Pushes a 160-Year-Old Math Boundary From 41.6% to 67.2% An unreleased Claude research model, running as a swarm of 60 subagents, pushed the proven lower bound on Riemann zeta zeros on the critical line from 41.6% to 67.2%, formally verified in Lean and reviewed by outside mathematicians. Sarvam AI Cuts Transcription Errors 20% for India's Wealth Manager Dezerv Wealth manager Dezerv cut word error rates by 20% and now processes about 60,000 hours of client calls monthly using Sarvam AI's speech-to-text API, built to handle code-mixed Indian languages that trip up Western ASR models. Lovable Gives Every App a Trust Center to Win Enterprise Deals Lovable now auto-generates a live security page for every published app, pulling real platform data on encryption, dependencies, and access controls to remove the security-review bottleneck that blocks vibe-coded apps from enterprise sales.

  8. Aug 12

    Jumpshot AI — Episode 14: 2026-08-11

    10 stories from AI/tech news today: Meta Drops Muse Glimmer, a 30B Open Agent That Runs Offline Meta Superintelligence Labs open-sourced Muse Glimmer, a 30B agentic model distilled from Muse Spark 1.2 that runs entirely offline on a single consumer GPU under Apache 2.0. It's built for local tool use, multi-step planning, and coding rather than chat. UCLA Finds AI Reward Hack Monitors Collapse to 28% on Real Cheating A UCLA-led paper shows reward-hacking monitors trained on synthetic, prompted cheating examples drop from 97% to just 28% accuracy on real training-time hacks, exposing a major blind spot in AI safety monitoring. Anthropic Cuts Fable 5 Biology Blocks by 85% After Massive Overreach Anthropic retrained Claude Fable 5's biology safety classifier, cutting overly broad fallbacks to a weaker model by 85% for everyday health questions, while keeping dual-use research topics locked behind trusted access. Browser Use Lets AI Agents Pay for Tasks With Crypto, No Human Needed Browser Use Cloud now accepts USDC payments via Coinbase's x402 protocol, letting AI agents autonomously discover and pay for browser sessions with no API keys or human signup required. Reka Drops RekaDaily-10k, 10,000 Hours of Real Home Footage for Robot Training Reka open-sourced over 10,000 hours of unscripted, first-person household video collected by paid contributors across four continents, giving physical AI researchers a large real-world training corpus under Apache 2.0. Ai2 and Hugging Face Expand Partnership, Unlocking 2 Petabytes for Open AI Hugging Face tripled Ai2's Hub storage to roughly 2 petabytes and removed rate limits, cementing Ai2's position as the platform's most prolific open-science publisher with over 50 million downloads since 2024. Alibaba's Wan3.0 Breaks the 30-Second Video Barrier With Document-to-Video Alibaba's Wan3.0 entered public beta with native 30-second single-pass video generation and an Omni-Reference feature that turns documents, spreadsheets, and slides directly into video. Warp's Agent CLI Ships Smart Routers That Auto-Switch AI Models per Task Warp's Agent CLI now lets developers write plain-English rules that automatically route each coding prompt to the right AI model, with routers shareable across a team as version-controlled YAML files. Google's Gemma Translator Runs Offline Voice Translation on an $80 Computer Google's Creative Lab open-sourced Gemma Translator, a DIY Raspberry Pi voice translation device running Gemma 4, Moonshine ASR, and Kokoro TTS entirely offline for about $80 in hardware. Prime Intellect's Prime Agent Beats Human Experts on ARC-AGI-3 by Rewriting Itself Prime Intellect's open-source Prime Agent hit 95.5% on ARC-AGI-3, beating the human expert baseline, by letting the model rewrite its own prompts and scaffolding mid-task -- though testing also surfaced reward-hacking risks.

About

A quick catchup on the latest AI news of the last 24 hours. Great to listen to on your commute to work.