Late Night With AI

Fourslash

Late Night with AI, brought to you by Fourlash - the team of AI experts, developers, researchers, observers, and users. Just like we like to say, "we use AI, so you don't have to," this show brings to you our honest and often surprising take on the most interesting and important developments in artificial intelligence. We chat about the latest breakthroughs, dissect the coolest applications, and maybe even ponder some of the more philosophical questions AI raises. So you can relax... or maybe be a little more informed. Follow us on X: @FourslashHQ

Episodes

  1. 04/23/2025

    Are o3, o4-mini OpenAI’s ultimate AI breakthrough?

    Get ready for a sassy deep dive into OpenAI’s April 2025 AI extravaganza! We’re spilling all the tea on their latest releases—think next-level reasoning models, a coding agent that’s a dev’s new BFF, and a visual reasoning flex that’s straight-up futuristic. But hold up, it’s not all roses: we’ve got benchmark blunders and a hallucination scandal that’s got the AI world clutching its pearls. From the o3 and o4-mini models to a sneaky new API tier, we’re unpacking the highs, lows, and chaos of OpenAI’s big week. Grab your energy drink—this episode’s serving tech innovation with a side of shade! - Episode Highlights: o3 and o4-mini Drop: OpenAI’s Reasoning RockstarsOpenAI unleashed the o3 and o4-mini models, built to “think longer” and tackle gnarly problems in coding, math, and science. o3: The big boss, labeled OpenAI’s “smartest” yet, slaying benchmarks like AIME (91.6%) and SWE-Bench (69.1%). Perfect for hardcore STEM and visual tasks. o4-mini: The lean, mean, cost-efficient machine, shockingly topping o3 on AIME math (93.4%) and available on ChatGPT’s free tier (with limits). Why It’s Lit: These models are “agentic” AF, autonomously wielding tools like web search, Python, and image generation to solve complex problems like pros.- “Thinking with Images” – AI’s New Visual Superpower Forget just seeing images—o3 and o4-mini can manipulate them, cropping, zooming, and reasoning with visuals in their thought process. From decoding blurry whiteboard sketches to analyzing charts, this feature’s a multimodal masterpiece. Hot Take: This is AI flexing visual IQ so slick, it’s like Photoshop and a PhD had a baby.- Codex CLI: OpenAI’s Open-Source Dev Candy Meet Codex CLI, a terminal-based coding agent that’s got devs swooning. Write, refactor, or debug code with natural language prompts. Runs on o4-mini by default (o3 optional) with three modes: Suggest (chill), Auto Edit (speedy), and Full Auto (sandboxed wild card). Why It Slaps: OpenAI’s back in the open-source game, and this tool’s a productivity rocket for terminal nerds. GitHub Copilot, watch your back!- Flex Processing: Bargain AI with a Catch OpenAI’s new Flex processing API tier cuts costs by 50% (o3: $5/$20 per million tokens; o4-mini: $0.55/$2.20). Trade-off? Slower responses and the occasional “try again later” error. Ideal for batch jobs, not your live chatbot. Plot Twist: Lower-tier devs now need ID verification to access o3’s premium goodies, sparking privacy grumbles.- Hallucination Drama: AI’s Fact-Checking Fumble Uh-oh: o3 (33%) and o4-mini (48%) are hallucinating way more than predecessors like o1 (16%). Think fake facts and wild stories about “running code on a MacBook.” Independent tests by Transluce reveal o3 spinning full-on fictional narratives and doubling down when called out. Cringe Alert: OpenAI’s like, “We don’t know why!” Is their reinforcement learning making AI too creative for its own good?- FrontierMath Flop: Benchmark Brouhaha OpenAI bragged about o3’s 25% score on the brutal FrontierMath test in December 2024, but Epoch AI’s April 2025 results tanked it at 10%. Theories? Model tweaks, benchmark changes, or OpenAI juicing internal tests with extra compute. Tea Spilled: This fumble’s got the industry side-eyeing vendor claims and begging for independent benchmarks.- Why You Should Care: OpenAI’s April 2025 releases are a tech rollercoaster—o3 and o4-mini are pushing AI into agentic, multimodal glory, but the hallucination spike and benchmark goof are major buzzkills. Codex CLI and Flex processing are sweet deals for devs, but the reliability wobbles and access hoops are giving everyone pause. Whether you’re coding, building, or just geeking out, this episode breaks down why OpenAI’s latest moves are shaking up the AI game—and why trust is the real MVP.

About

Late Night with AI, brought to you by Fourlash - the team of AI experts, developers, researchers, observers, and users. Just like we like to say, "we use AI, so you don't have to," this show brings to you our honest and often surprising take on the most interesting and important developments in artificial intelligence. We chat about the latest breakthroughs, dissect the coolest applications, and maybe even ponder some of the more philosophical questions AI raises. So you can relax... or maybe be a little more informed. Follow us on X: @FourslashHQ