AI Explained

AI Explained

Covering the biggest news of the century - the arrival of smarter-than-human AI. The author of Simple Bench, exposing the remaining human-LLM reasoning gap. Solo Developer of LM Council: https://lmcouncil.ai - Code - INSIDER15Join me at AI Insiders, with exclusive videos and a 1000+ network of AI enthusiasts and professionals: https://www.patreon.com/AIExplainedLinks:AI Insiders: Code - INSIDER15: https://www.patreon.com/AIExplainedMy AI Council App: https://lmcouncil.aiSimple Bench: https://simple-bench.com/Podcast (New!): https://aiexplainedopodcast.buzzsprout.com/Newsletter: https://signaltonoise.beehiiv.com/X: https://twitter.com/AIExplainedYT

  1. 25m ago

    AI is getting a little out of control

    Wow. Mathematical breakthroughs that would be called genius if done by humans. A secret message-board w/ AI agent swarms leaving notes read by future versions. Hassabis leaves CEO position, or was pushed out? Not to mention news of constitutional breakdowns, Gemini 4 and Jeff Dean…https://80000hours.org/aiexplainedExclusive Videos - AI Insiders ($7/month if annual!): https://www.patreon.com/AIExplainedChapters:00:00 - Introduction01:16 - 10 Autonomous Discoveries08:20 - The Security ‘Incident’15:29 - MessageBoard19:50 - Constitutional Failure24:30 - Google Explosion29:12 - Closing ThoughtsSecurity Incident: Paper: https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/6a724858f7db25c81487016d_Security%20Incident%20INC-2026-07-28-01.pdfPost: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testinghttps://www.theguardian.com/technology/2026/aug/05/ai-models-have-been-going-rogue-in-tests-how-worried-should-we-beMeta too: https://www.theinformation.com/articles/meta-ai-model-hacked-another-company-cybersecurity-testing?rc=sy0ihqBlatantly Misaligned: https://x.com/yonashav/status/2085167279893795022No Excuses: https://x.com/boazbaraktcs/status/2085034783541964945Surreal Moment: https://x.com/mobav0/status/2084341687883841732Wired Article: https://archive.is/20260806002210/https://www.wired.com/story/openai-didnt-notice-its-ai-agents-using-a-message-board-to-plan-their-hacking-spree/Chunky Post-training: https://x.com/johnschulman2/status/2084835800899076313Watershed Moment: https://www.groundlevel-ai.com/p/openai-gives-first-detailed-debrief?has_completed_unsubscribed_unlock=trueHedgeFund Hack: https://finance.yahoo.com/technology/ai/articles/major-hedge-funds-targeted-wave-154044981.html10 Discoveries:Paper: https://cdn.openai.com/pdf/ten-proofs-oai.pdfPost: https://openai.com/index/ten-advances-in-mathematics/Haven’t Solved Math: https://x.com/polynoamial/status/2083476852216369294Half with Fable: https://x.com/__alpoge__/status/2083855298239078748Pivot to Safety: https://www.understandingai.org/p/mathematicians-are-grappling-withAmodei Essay: https://darioamodei.com/essay/the-adolescence-of-technology?utm_source=chatgpt.comConstitution: https://www.anthropic.com/constitutionMidtraining: https://arxiv.org/pdf/2605.02087DroneBench: https://andonlabs.com/evals/drone-benchhttps://x.com/andonlabs/status/2085125235188310445Book Deal: https://x.com/venturetwins/status/2085185278378222054Making Marble: https://x.com/Rainmaker1973/status/2084560915404382685Google News:Hassabis Move: https://blog.google/company-news/inside-google/message-ceo/next-chapter-ai-momentum/Resignation: https://x.com/Turn_Trout/status/2077448610157891734Periodic Labs: https://periodic.com/Jeff Dean: https://x.com/JeffDean/status/208503460417260372414 Challenges: https://gcsp.engineering.asu.edu/apply/become-a-grand-challenge-scholar/the-14-grand-challenges-for-engineering/Going Places for Sure: https://x.com/thsottiaux/status/2085223189555126579Gemini 4: https://x.com/firstadopter/status/2085215060532535449Pacing Frontier Patreon Video: https://www.patreon.com/AIExplained/posts/opus-5-amodei-165170363Non-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

    AI is getting a little out of control
  2. 1h ago

    What the Freakiness of 2025 in AI Tells Us About 2026

    It’s probably not possible to satisfactorily condense a 12 month’s worth of weird progress in AI, as well as predictions for the year to come, into one video. But I’m gonna try anyway because it has been a very strange time.http://matsprogram.org/s26-aieMy new app! https://lmcouncil.aiPatreon Interview: https://www.patreon.com/posts/robot-in-your-27-146376094Chapters:00:00 - Introduction00:34 - Reasoning Models … and limits02:54 - A playable world03:36 - Realism03:50 - AI Slop gone mainstream05:03 - DolphinGemma05:39 - Public Mood07:34 - AI Enlisted08:30 - GPT-511:05 - Open Weight not out13:00 - METR Breakout17:30 - VASA-118:28 - Lateral Productivity20:15 - 1 or 1000 benchmarks needed?24:54 - Continual Learning + Altman on Superintelligence28:08 - Automated Information Discovery ft AlphaEvolveHassabis on Generality: https://x.com/demishassabis/status/2003097405026193809https://www.youtube.com/watch?v=PqVbypvxDtoGemini 3: https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini_3_table_final_HLE_Tools_on.gifReasoning Trade-offs: https://arxiv.org/pdf/2504.13837DolphinGemma: https://blog.google/technology/ai/dolphingemma/?s=09Genie 3: https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models/METR Time Horizon: https://arxiv.org/pdf/2503.14499https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/Flaws: https://x.com/ShashwatGoel7/status/2002369517499105443https://shash42.substack.com/p/how-to-game-the-metr-plothttps://x.com/METR_Evals/status/2002203627377574113GPT-5 - Altman phd in everything: https://edition.cnn.com/2025/08/14/business/chatgpt-rollout-problemshttps://simple-bench.com/AI Slop: https://www.youtube.com/watch?v=I_3vxoJDD9khttps://www.theguardian.com/technology/2025/dec/16/boost-for-artists-in-ai-copyright-battle-as-only-3-per-cent-back-uk-active-opt-out-planSurvey: https://x.com/SearchlightInst/status/2001057144842387920/photo/1Nvidia Nemotron: https://x.com/percyliang/status/2000608134205985169OpenAI Compute Flywheel: https://x.com/OpenAI/status/2001363007209914399/photo/1Altman Interview: https://www.youtube.com/watch?v=2P27Ef-LLuQAI in Govt: https://x.com/jdcmedlock/status/1939814516503847259Benchmark Gaming: https://techcrunch.com/2025/04/07/meta-exec-denies-the-company-artificially-boosted-llama-4s-benchmark-scores/AlphaEvolve: https://deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/AlphaEvolve.pdf?utm_source=deepmind.google&utm_medium=referral&utm_campaign=gdm&utm_content=Continual Learning: https://abehrouz.github.io/files/NL.pdfJob Risk: https://archive.ph/20250708204527/https://www.axios.com/2025/05/28/ai-jobs-white-collar-unemployment-anthropicGPT4o: https://x.com/AISafetyMemes/status/1916889492172013989Vasa-1: https://www.microsoft.com/en-us/research/project/vasa-1/Three Views: https://www.lesswrong.com/posts/K2D45BNxnZjdpSX2j/ai-timelinesTuring Test: https://x.com/tunguz/status/1907185471211422147Karpathy Year in Review: https://karpathy.bearblog.dev/year-in-review-2025/LLM Brainrot: https://arxiv.org/pdf/2510.13928Lateral Productivity: https://www.aisi.gov.uk/frontier-ai-trends-reportEmotional Quotient: https://arxiv.org/pdf/2511.08394Non-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/AI Insiders ($9!): https://www.patreon.com/AIExplainedReasoning Models would boost results but not change paradigmGenie 3 makes the world playable (literally, in the case of yesterday’s news)DolphinDecodingVeo 3.1 / Sora 2 / Nano Banana Pro / Elevenlabs Voice/MusicBut AI slop everywherePublic attitude very mixed (Hassabis)AI gets e Learn more about your ad choices. Visit megaphone.fm/adchoices

    What the Freakiness of 2025 in AI Tells Us About 2026
  3. 1h ago

    Anthropic: Our AI just created a tool that can ‘automate all white collar work’, Me:

    A new tool, with code written *only* by AI, has gone omega-viral: Claude Cowork. But is the hype justified? What do the stats say on productivity? Where is the truth in a sea of noise? What is truth? Can we handle the truth? Where's Nemo?https://matsprogram.org/s26-aieCheck out my new app! https://lmcouncil.aiAI Insiders ($9!): https://www.patreon.com/AIExplainedChapters: 00:00 - Introduction01:12 - Claude Cowork07:36 - Productivity Speed-up + jobs10:19 - Comparing Models12:46 - Brittle AI PaperCowork Intro: https://x.com/claudeai/thread/2010805682434666759'All of it': https://x.com/bcherny/status/2010813886052581538'AGI' Claims: https://x.com/deepfates/status/2004994698335879383Douglas Interview: https://www.youtube.com/watch?v=TOsNrV3bXtQ&t=2313sJob Stats: https://www.oxfordeconomics.com/wp-content/uploads/2026/01/Evidence-of-an-AI-driven-shakeup-of-job-markets-is-patchy.pdfAmodei Prediction: https://fortune.com/2025/05/28/anthropic-ceo-warning-ai-job-loss/GenAI Traffic: https://x.com/demishassabis/status/2009075877347512545Illusion of Insight: https://arxiv.org/pdf/2601.00514Entropy Exploration: https://arxiv.org/pdf/2506.14758ProRL: https://arxiv.org/pdf/2505.24864Genesis Mission: https://www.whitehouse.gov/presidential-actions/2025/11/launching-the-genesis-mission/https://deepmind.google/blog/how-were-supporting-better-tropical-cyclone-prediction-with-ai/Non-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

    Anthropic: Our AI just created a tool that can ‘automate all white collar work’, Me:
  4. 2h ago

    ChatGPT Fails Basic Logic but Now Has Vision, Wins at Chess and Prompts a Masterpiece

    ChatGPT will now have vision, but can it do basic logic? I cover the latest news - including GPT Chess! - as well as go through almost a dozen papers and how they relate to the central question of LLM logic and rationality. Starring the Reversal Curse and featuring conversations with two of the authors at the heart of it all. I also get to a DALL-E 3 vs Midjourney comparison, MuZero, MathGLM, Situational Awareness and much more!https://www.patreon.com/AIExplainedOpenAI GPT-V (Hear and Speak): https://openai.com/blog/chatgpt-can-now-see-hear-and-speakReversal Curse: https://owainevans.github.io/reversal_curse.pdfMahesh Tweet: https://twitter.com/madiator/status/1705376797293183208Neel Nanda Explanation: https://twitter.com/NeelNanda5/status/1705995593657762199Karpathy tweet: https://twitter.com/karpathy/status/1705322159588208782Trask Explanation: https://twitter.com/iamtrask/status/1705361947141472528Play Chess vs GPT 3.5 Instruct: https://parrotchess.com/Paige Bailey on Cognitive Revolution: https://www.youtube.com/watch?v=K-XYxLifpQEAvenging Polanyi's Revenge: https://m-cacm.acm.org/magazines/2021/2/250077-polanyis-revenge-and-ais-new-romance-with-tacit-knowledge/abstractFaith and Fate Paper: https://arxiv.org/pdf/2305.18654.pdfCounterfactuals Paper: https://arxiv.org/pdf/2307.02477.pdfLesswrong AGI Timelines: https://www.lesswrong.com/posts/SCqDipWAhZ49JNdmL/paper-llms-trained-on-a-is-b-fail-to-learn-b-is-a?commentId=bkxcTqAtYW8wgHKb5Professor Rao Paper w/ Blocksworld: https://arxiv.org/pdf/2305.15771.pdfMath Based on Number Reasoning: https://aclanthology.org/2022.findings-emnlp.59.pdfMuZero: https://www.deepmind.com/blog/muzero-mastering-go-chess-shogi-and-atari-without-ruleshttps://www.nature.com/articles/s41586-020-03051-4.epdf?sharing_token=kTk-xTZpQOF8Ym8nTQK6EdRgN0jAjWel9jnR3ZoTv0PMSWGj38iNIyNOw_ooNp2BvzZ4nIcedo7GEXD7UmLqb0M_V_fop31mMY9VBBLNmGbm0K9jETKkZnJ9SgJ8Rwhp3ySvLuTcUr888puIYbngQ0fiMf45ZGDAQ7fUI66-u7Y%3DEfficient Zero: https://arxiv.org/pdf/2111.00210.pdfLet’s Verify Step by Step OpenAI paper: https://cdn.openai.com/improving-mathematical-reasoning-with-process-supervision/Lets_Verify_Step_by_Step.pdfMy Video on That: https://www.youtube.com/watch?v=hZTZYffRsKI&t=5sSuperintelligence Poll: https://www.vox.com/future-perfect/2023/9/19/23879648/americans-artificial-general-intelligence-ai-policy-poll?s=09Anthropic Announcement: https://www.anthropic.com/index/anthropic-amazonDALL-E 3 Tweet Thread: https://twitter.com/OfficialLoganK/status/1704850313889595399 https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

    ChatGPT Fails Basic Logic but Now Has Vision, Wins at Chess and Prompts a Masterpiece
  5. 3h ago

    Enter PaLM 2 (New Bard): Full Breakdown - 92 Pages Read and Gemini Before GPT 5? Google I/O

    Google puts it foot on the accelerator, casting aside safety concerns to not only release a GPT 4 -competitive model, PaLM 2, but also announce that they are already training Gemini, a GPT 5 competitor [likely on TPU v5 chips]. This is truly a major day in AI history, and I try to cover it all. I'll show the benchmarks in which PaLM (which now powers Bard) beats GPT 4, and detail how they use SmartGPT-like techniques to boost performance. Crazily enough, PaLM 2 beats even Google Translate, due in large part to the text it was trained on. We'll talk coding in Bard, translation, MMLU, Big Bench, and much more.I'll end on the Universal Translator deepfakes and the underwhelming results from Sundar Pichai and Sam Altman's trip to the White House and what Hinton says about it all. On a more positive note, I cover Med PaLM 2, which could genuinely save thousands of lives. PaLM 2 Technical Report: https://ai.google/static/documents/palm2techreport.pdfRelease Notes Google Blog: https://blog.google/technology/ai/google-palm-2-ai-large-language-model/Bard Access: https://bard.google.com/Scaling Transformer to 1M tokens: https://arxiv.org/pdf/2304.11062.pdfGPT 4 Technical Report: https://arxiv.org/pdf/2303.08774.pdfBard Languages: https://support.google.com/bard/answer/13575153?hl=enSelf Consistency Paper: https://arxiv.org/pdf/2203.11171.pdfAre Emergent Abilities a Mirage: https://arxiv.org/pdf/2304.15004.pdfSparks of AGI Paper: https://arxiv.org/pdf/2303.12712.pdfBig Bench Hard: https://github.com/suzgunmirac/BIG-Bench-HardGoogle Keynote: https://www.youtube.com/watch?v=cNfINi5CNbYGemini: https://www.youtube.com/watch?v=1UvUjTaJRz0Med PaLM 2: https://www.youtube.com/watch?v=k_-Z_TkHMqATPU v5: https://ai.googleblog.com/2022/01/google-research-themes-from-2021-and.htmlHinton Warning: https://www.youtube.com/watch?v=FAbsoxQtUwMWhite House Readout: https://www.whitehouse.gov/briefing-room/statements-releases/2023/05/04/readout-of-white-house-meeting-with-ceos-on-advancing-responsible-artificial-intelligence-innovation/https://www.patreon.com/AIExplained Non-Hype, Free Newsletter: https://signaltonoise.beehiiv.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

    Enter PaLM 2 (New Bard): Full Breakdown  - 92 Pages Read and Gemini Before GPT 5? Google I/O
  6. 4h ago

    OpenAI Tests if GPT-5 Can Automate Your Job - 4 Unexpected Findings

    An OpenAI report released in the last 24 hours is the best look we have as to whether 2025 AI can automate your job. I’ll go through 4 unexpected findings, from which model is best at what, to practical tips and massive caveats. Plus UFC robots, radiologist essay, don’t trust videos and the blockers to the singularity. Gray Swan: https://app.grayswan.ai/ai-explainedAI Insiders ($9!): https://www.patreon.com/AIExplainedChapters:00:00 - Introduction00:55 - OpenAI Report Summary02:40 - Tipping Point Speed-up04:11 - Better than Industry Experts?06:33 - Big Caveat11:10 - Karpathy and the Radiologist Analogy13:30 - OutroGDPval: https://cdn.openai.com/pdf/d5eb7428-c4e9-4a33-bd86-86dd4bcf12ce/GDPval.pdf[GDP Impact: https://fred.stlouisfed.org/release/tables?rid=331&eid=211Task List: https://www.onetonline.org/link/summary/11-9141.00Summer Tweet: https://x.com/LHSummers/status/1971252567981146347Emad: https://x.com/EMostaque/status/1971254153067593739Robots: https://x.com/cixliv/status/1967663286679478759Unitree G1: https://x.com/UnitreeRobotics/status/1970039940022239491Don’t Trust Video: https://x.com/AISafetyMemes/status/1970453369446871420AGI Tweet: https://x.com/hyhieu226/status/1968378785709133915Blockers to the Singularity: https://www.patreon.com/posts/blockers-to-and-139264812Framework: https://gemini.google.com/share/f4b9c85a6ae9METR Study (Dev Slowdown): https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/Karpathy Tweet: https://x.com/karpathy/status/1971220449515516391Radiology Essay: https://worksinprogress.co/issue/the-algorithm-will-see-you-now/Non-hype Newsletter: https://signaltonoise.beehiiv.com/Podcast: https://aiexplainedopodcast.buzzsprout.com/ Learn more about your ad choices. Visit megaphone.fm/adchoices

    OpenAI Tests if GPT-5 Can Automate Your Job - 4 Unexpected Findings

About

Covering the biggest news of the century - the arrival of smarter-than-human AI. The author of Simple Bench, exposing the remaining human-LLM reasoning gap. Solo Developer of LM Council: https://lmcouncil.ai - Code - INSIDER15Join me at AI Insiders, with exclusive videos and a 1000+ network of AI enthusiasts and professionals: https://www.patreon.com/AIExplainedLinks:AI Insiders: Code - INSIDER15: https://www.patreon.com/AIExplainedMy AI Council App: https://lmcouncil.aiSimple Bench: https://simple-bench.com/Podcast (New!): https://aiexplainedopodcast.buzzsprout.com/Newsletter: https://signaltonoise.beehiiv.com/X: https://twitter.com/AIExplainedYT