System Prompt

Peter

System Prompt is a podcast about what’s actually happening in AI. Not hype. Not surface-level takes. We break down how AI is changing software, SaaS, infrastructure, and the way systems are built focusing on real-world tradeoffs, architecture decisions, and where the value is actually shifting. If you’re building, deploying, or thinking seriously about AI, this is for you.

  1. 5d ago

    You Bought AI Hardware. Now What? | vLLM Setup Explained

    READ THE FULL EPISODE PAGE :https://devmesh.tech/podcast/youve-bought-ai-hardware-now-what Buying AI hardware is only the beginning. In Episode 28 of System Prompt, Peter and Val follow up on Before You Buy AI Hardware, Watch This by moving from planning into the actual hosting layer. Using GLM-5.3-Flash across two NVIDIA DGX Sparks with vLLM, they break down what the serving command is doing, which settings actually matter, and why hosting a model is very different from simply loading one. The conversation covers tensor parallelism, distributed execution, RoCE and NCCL networking, Docker, model mounts, KV cache, context length, concurrency, Mixture-of-Experts execution, chat templates, reasoning and tool-call parsers, DFlash speculative decoding, prefill, decode, and live serving metrics.Then they start the real cluster, bring the worker online before the head node, watch the model load across both Sparks, and run GLM-5.3-Flash through a coding workload. The live metrics show speculative decoding acceptance changing with the task, with generation throughput moving from the teens into the 30s and 40s as the workload becomes more predictable. • What happens after you buy AI hardware • Hosting GLM-5.3-Flash across two DGX Sparks • vLLM and distributed inference • Tensor parallelism across multiple devices • RoCE, NCCL, and cluster communication • Docker, model mounts, and runtime configuration • Context windows and KV cache • FP8 KV cache vs. FP16 • Concurrency and max active sequences • Mixture-of-Experts execution • Chat templates, reasoning parsers, and tool-call parsers • DFlash speculative decoding • Prefill, decode, and generation throughput • Why real workloads have to be tested QUICK CORRECTION At one point I say the ConnectX-7 link is 200 gigabytes per second. I meant 200 gigabits per second, or 200 Gb/s. Yes, I know the difference. It was a live recording, I misspoke, and I am correcting it here. KEY TAKEAWAYS RUNNING A MODEL IS NOT THE SAME AS HOSTING ONE Loading the weights is step one. Hosting adds networking, scheduling, memory management, concurrency, cache behavior, runtime compatibility, and observability. THE MODEL IS ONLY ONE PART OF THE STACK Runtime, quantization, KV cache, speculative decoding, context limits, scheduling, and hardware topology all affect system behavior. YOU DO NOT NEED TO ENGINEER EVERY LAYER YOURSELF Start with a known-good path, understand the major knobs, test it on your hardware, and go deeper only where the workload proves you need to. SPECULATIVE DECODING ONLY HELPS WHEN THE DRAFT IS RIGHT When DFlash acceptance is low, speculative work can become overhead. When the workload becomes more predictable, acceptance and generation throughput can rise significantly. THE WORKLOAD DETERMINES THE CONFIGURATION There is no universal best setup. Context, concurrency, precision, cache size, and speculative decoding all need to be tested against the actual task. CHAPTERS 00:00 Intro and the Follow-Up 03:23 Two DGX Sparks, vLLM, and GLM-5.3-Flash 07:00 Breaking Down the Hosting Command 11:00 Model Mounts, DFlash, and Runtime Patches 14:12 NCCL, RoCE, and Distributed Networking 17:30 API Behavior, Tool Calls, Reasoning, and Chat Templates 21:40 The vLLM Knobs You Actually Tune 25:45 FP8 KV Cache and Memory Tradeoffs 27:30 DFlash and Speculative Decoding 31:38 Starting the Two-Node Cluster 41:51 Running GLM-5.3-Flash 45:56 Reading the Live Metrics 50:17 Why You Have to Test 54:22 Coding Workloads and Higher Draft Acceptance 59:37 What This Means for Businesses 1:04:01 Closing Thoughts

  2. Sep 25

    Before You Buy AI Hardware, Watch This

    READ THE FULL EPISODE PAGE https://devmesh.tech/podcast/before-you-buy-ai-hardware-watch-thisMost businesses do not need more AI hardware. They need a better understanding of the workload they are actually trying to run. In Episode 27 of System Prompt, Peter and Val break down inference runtime, the layer that sits between the model and the hardware serving it, and why runtime decisions matter before a business starts buying GPUs, DGX systems, or other local AI infrastructure. The conversation covers Ollama, llama.cpp, vLLM, SGLang, quantization, context windows, KV cache, prefill, decoding, batching, speculative decoding, and parallelism. But the bigger question is business capacity. How many users will actually be active at once? Which departments need the most inference? When does demand spike? What should run locally, and what should stay in the cloud? Using examples from a 150-person company to enterprise GPU deployments, Peter and Val argue that infrastructure should be sized from real usage data, not assumptions. A business may not need 150 concurrent sessions just because it has 150 employees, and it may not need more hardware when better scheduling can solve the problem. • What inference runtime actually is • Ollama, llama.cpp, vLLM, and SGLang • Why concurrency should be measured instead of guessed • How quantization affects memory, performance, and quality • Context windows, KV cache, prefill, and time to first token • Why continuous batching matters • Speculative decoding and runtime optimization • Local devices vs. shared inference • Why expensive hardware can sit mostly idle • How businesses should pilot before scaling KEY TAKEAWAYS AN UNTESTED CONFIGURATION IS WINDOW DRESSING If a model, runtime, quantization, and hardware configuration has not been tested against the real workflow, it is still an assumption. Production decisions should be based on measured behavior. HEADCOUNT IS NOT CONCURRENCY A 150-person company does not automatically need infrastructure for 150 simultaneous sessions. Real usage depends on departments, workflows, working hours, and demand patterns. RUNTIME CHANGES HOW HARDWARE GETS USED Batching, caching, scheduling, memory management, and parallelism all affect how efficiently requests turn into useful GPU work. START SMALL AND COLLECT REAL DATA Pilot with one department or workload, measure throughput and bottlenecks, then scale from evidence instead of buying for theoretical maximums. IDLE CAPACITY IS STILL WASTED MONEY A powerful GPU cluster does not create ROI by existing. If expensive accelerators sit idle most of the day, the infrastructure may be oversized or poorly scheduled. THE BUSINESS PROBLEM COMES FIRST Before choosing a model, runtime, or piece of hardware, define what the system needs to accomplish, how often it will be used, and what level of performance the workflow actually requires. CHAPTERS 00:00 Intro and 20,000 Views 01:51 Why Businesses Are Exploring Local AI 03:25 What Inference Runtime Actually Is 06:51 Real Concurrency vs. Headcount 10:29 llama.cpp, GGUF, and Quantization 14:13 An Untested Configuration Is Window Dressing 15:22 Context and KV Cache 17:06 vLLM and Continuous Batching 18:28 Prefill, Decode, and Time to First Token 21:12 Prefix Caching and Runtime Optimization 25:17 Speculative Decoding 29:05 Engineering for Capacity 30:47 Local Devices vs. Shared Infrastructure 35:09 When Local Inference Makes Sense 39:53 Idle Hardware and Wasted Spend 48:10 Scheduling Inference by Department 49:40 The Full AI Infrastructure Stack

  3. Sep 18

    Your AI Isn’t Dumb. It’s Missing Context.

    READ THE FULL EPISODE PAGE https://devmesh.tech/podcast/your-ai-isnt-dumb-its-missing-context Most people do not have an AI problem. They have a context problem. In Episode 26 of System Prompt, Peter and Val break down why AI gives generic answers when it is handed generic information, and how a structured context document can turn the same model into something much more useful for real business work. Using a cloud-services account example, they compare three approaches: a short paragraph with a vague question, the same paragraph with a better prompt, and a full client context profile built from account, usage, support, contract, and relationship data. The difference is not subtle. The model goes from plausible advice to specific, evidence-backed next steps tied to the actual customer. The conversation also gets into where that context should come from, how teams can combine CRM data, tickets, emails, usage metrics, notes, and subject-matter expertise, why humans still need to verify the final document, and how this can replace long cross-functional information-gathering meetings with a much tighter workflow. • Why AI gives generic answers even when the model is good • What a business context document actually is • Why better prompting helps — but only up to a point • How CRM, ticketing, email, usage, and contract data fit together • How structured context changes the quality of account recommendations • Why every risk and recommendation should be tied to evidence • How context documents can become a living source of truth • Why account executives and technical teams still need to verify the data • How prompts encode what your business actually cares about • Why context engineering matters more than performative prompting • How AI can reduce cross-functional meeting overhead • Why many AI opportunities are process-shaped business problems KEY TAKEAWAYS BETTER MODELS DO NOT FIX MISSING CONTEXT A strong model can still only reason from what you give it. If the information is thin, ambiguous, or disconnected, the output will usually be generic. PROMPTING HELPS, BUT CONTEXT CHANGES THE GAME A better prompt improves structure and discipline, but a full context document gives the model enough evidence to produce specific risks, opportunities, questions, and next steps. THE CONTEXT DOCUMENT SHOULD BE A LIVING SOURCE OF TRUTH Customer history, usage, support issues, contracts, upcoming changes, relationship details, and decisions should evolve with the account instead of being rebuilt from scratch every time. HUMANS STILL OWN THE FACTS AND THE DECISIONS AI can synthesize information, but the people closest to the account still need to confirm the numbers, correct drift, and decide what actually matters. THE REAL ROI IS PROCESS COMPRESSION The value is not that AI “takes over” the work. It can compress hours of gathering, reconciling, and summarizing information into a much shorter review process. CHAPTERS 00:00 Why Context Matters 02:00 What a Context Document Is 07:59 The Generic Account Example 12:48 Better Prompt, Better Output 16:40 Building the Full Client Context Profile 21:01 What Changes With Complete Context 26:19 Replacing Long Cross-Team Meetings 29:37 Context + Prompt = Business Logic 32:12 Process Problems, Not Just Technical Problems 33:52 Guardrails, Verification, and Workflow Design

  4. Sep 9

    The Reality of Physical AI w/ Hiten Sonpal

    READ THE FULL EPISODE PAGE https://devmesh.tech/podcast/the-reality-of-physical-ai-with-hiten-sonpal Physical AI gets a lot more serious when a bad answer can move thousands of pounds of machinery. In Episode 25 of System Prompt, Peter and Val sit down with Hiten Sonpal, CEO of RISE Robotics and a robotics veteran whose career spans iRobot, defense robotics, autonomous lawn care, Electric Sheep, and heavy industrial machinery. The conversation gets into what changes when probabilistic AI meets deterministic machines, why safety envelopes matter, and why Hiten believes the future of robotics is more likely to be specialized machines that amplify people than general-purpose humanoid replacements. Hiten also breaks down RISE Robotics’ Beltdraulic technology, the SuperJammer robotic arm that lifted more than 7,000 pounds, where AI is actually useful in engineering today, and why impressive technology still does not become a business until customers want it and will pay for it. • What changes when AI starts moving physical machines • Human-in-the-loop robotics and deterministic safety systems • Why AI should not have unrestricted control of heavy equipment • RISE Robotics’ Beltdraulic alternative to hydraulics • The SuperJammer arm and its 7,000+ pound world-record lift • What Beltdraulics does better — and worse — than hydraulics • How RISE uses AI for software, sourcing, and calculations • Why robotics has a harder training-data problem than language models • Desirability, viability, and feasibility as tests for real innovation • Why Hiten is skeptical of humanoid-robot hype • Specialized robots versus general-purpose human replacements • What software engineers misunderstand about hardware • What robotics engineers underestimate about modern AI • Which jobs robotics is most likely to automate first KEY TAKEAWAYS PHYSICAL AI NEEDS DETERMINISTIC BOUNDARIES A model can make high-level decisions, but safety-critical machinery still needs hard constraints. The system has to know when the answer is simply “no.” SPECIALIZATION BEATS HUMAN REPLACEMENT A robot does not need arms, legs, hands, and twenty degrees of freedom if the task can be solved with wheels and a purpose-built mechanism. Complexity has to earn its place. HARDWARE CHANGES THE FAILURE MODEL Software bugs can often be patched quickly. Hardware failures involve parts, lead times, prototypes, shipping, installation, and sometimes redesigning the machine itself. NOVELTY IS NOT A BUSINESS A technically impressive product still has to solve a problem customers care about, at a price they will pay, with economics that work. AI IS ALREADY HELPING ENGINEERING — WITH CHECKS RISE is using AI to accelerate software work, research components, and assist with calculations, but Hiten is clear that engineering outputs still need human verification. CHAPTERS 00:00 Physical AI Gets Real 02:21 What Is Different About Robotics Now? 03:18 Failure, Safety, and Human Control 06:22 Should AI Directly Control Heavy Machinery? 09:03 Robots, Jobs, and Human Amplification 12:11 Why Robotics Needs a Hardware Revolution 13:05 The Origin of Beltdraulics 14:46 What Beltdraulics Does Worse Than Hydraulics 18:05 Building a 7,000-Pound-Lift Robotic Arm 20:42 How RISE Uses AI 24:11 The Robotics Training-Data Problem 27:01 Novel Technology vs. Real Business 28:52 The Problem With Humanoid Robots 36:06 The iRobot Lawn-Mower Story 39:58 Which Jobs Robotics Takes First 41:57 What AI People Misunderstand About Robotics 44:26 What Robotics Engineers Underestimate About AI 46:04 The Robotics Problem That Still Is Not Solved

  5. Sep 3

    NVIDIA, Asana, and New Models

    READ THE FULL EPISODE PAGE https://devmesh.tech/podcast/nvidia-asana-and-new-models NVIDIA agreeing to acquire Hugging Face changes more than who owns the biggest model repository in open-source AI. It may also make the open ecosystem harder to push around. In Episode 24 of System Prompt, Peter and Val break down NVIDIA’s $12.9 billion Hugging Face deal, Asana’s claim that Codex compressed years of engineering work into weeks, and a wave of open models getting harder to dismiss. The conversation moves past the headlines into the systems underneath them: why NVIDIA benefits from open models thriving, why AI productivity claims need methodology and labor context, and why better local models are changing how developers think about cost, privacy, and infrastructure. Peter also walks through a one-shot AI Signal dashboard built with GLM 5.3 Flash and explains why improving open models are forcing him to rethink how much scaffolding smaller models need. WHAT WE DISCUSS • NVIDIA’s $12.9 billion Hugging Face acquisition • Why NVIDIA may become a powerful defender of open-source AI • Whether Hugging Face can stay neutral across CUDA, AMD, MLX, and other runtimes • Why open models may become harder to marginalize • AI-assisted hacking, agency, and tool access • Anthropic’s pricing problem as open models improve • Why model quality alone may not protect a frontier-model moat • Asana compressing years of engineering work into weeks with Codex • Why AI productivity claims need labor and methodology context • GLM 5.3 Flash and the rise of stronger open-weight models • Building an AI news dashboard from one ambiguous prompt • Qwen, GLM, and model routing for planning, execution, and review • Why assumptions about smaller models may already be outdated KEY TAKEAWAYS NVIDIA CHANGES THE OPEN-SOURCE POWER BALANCE Lobbying against a smaller open-model company is one thing. Doing it when NVIDIA has billions invested in the ecosystem is another. If open models grow, NVIDIA is positioned to benefit. HEADLINES NEED SYSTEM CONTEXT Compressing five years of work into two weeks sounds incredible, but the model is only part of the system. Domain expertise, testing, review, infrastructure, labor, and methodology shape the result. MODEL MOATS ARE GETTING THINNER Anthropic still produces excellent models, but quality is no longer the only variable. Cost, usage limits, privacy, local deployment, and capable open models all affect where workloads go. OPEN MODELS NEED LESS HAND-HOLDING Smaller models once needed heavy scaffolding to produce reliable results. That assumption is eroding as newer models improve at reasoning, coding, review, and ambiguous product decisions. ARCHITECTURE HAS TO FOLLOW THE MODELS Systems built around yesterday’s models may not make sense for tomorrow’s. Better models change where guardrails belong, which tasks need frontier models, and how much infrastructure is necessary. CHAPTERS 00:00 NVIDIA Buys Hugging Face 05:47 Does NVIDIA Protect Open Source? 11:34 Open Models, Hacking, and Agency 18:17 Anthropic’s Shrinking Moat 22:04 The Asana Codex Story 27:16 What AI Productivity Headlines Leave Out 31:13 GLM 5.3 Flash Arrives 34:00 Building AI Signal with an Open Model 40:50 One Prompt, Better Product Decisions 43:42 Price, Privacy, and Frontier Models 49:44 Rethinking Smaller Open Models

  6. Aug 27

    AI, Accountants, and the Future of the Profession

    READ THE FULL EPISODE PAGE https://devmesh.tech/podcast/ai-accountants-and-the-future-of-the-profession AI is not just changing the tools accountants use. It is changing the work itself. In Episode 23 of System Prompt, Peter and Val are joined by Peter McCarroll, founder of Fuel Accountants and creator of The AI Accountant, to talk about what AI means for accounting firms, their teams, and the future of the profession. The conversation moves past prompt training and AI hype into the harder questions: which work should be automated, what still requires human judgment, how firms should handle sensitive client data, and what happens when AI starts taking over work that used to require years of experience. Peter McCarroll also shares a real AI failure that cost his firm roughly 10 hours of rework, why accountants should train on the work instead of the tools, and why he believes firms may have only a couple of busy seasons left before the profession looks very different. WHAT WE DISCUSS • Why AI training often fails to change how people work • Training on workflows instead of individual AI tools • Why the billable hour may become a weaker measure of productivity • Where automation should stop in professional services • Human accountability for AI-generated work • What happens when AI gets accounting work wrong • Using deterministic tools like Python for calculations • Choosing between Claude, ChatGPT, Gemini, and open models • Token cost and model selection for agentic workflows • Protecting sensitive client and financial data • Shadow AI, retention, and business risk • Whether AI will reduce accounting jobs • Why professional judgment may not be the moat accountants think it is • Context engineering for client-specific financial analysis • How accounting moves from the rearview mirror to the dashboard KEY TAKEAWAYS TRAIN ON THE WORK, NOT THE TOOL Teaching someone how to use ChatGPT or Claude does not automatically change a workflow. Start with the work, redesign the process, then teach the team how AI fits into it. ACCOUNTABILITY DOES NOT GET AUTOMATED AI can perform more of the work, but the accountant still has to stand behind what reaches the client. Professional services cannot outsource responsibility to a model. AI NEEDS CHECKPOINTS A model can appear to follow instructions while quietly dropping part of the work. Verification, review, and deterministic steps matter when the output affects financial decisions. PROFESSIONAL JUDGMENT IS CHANGING Much of what experts call judgment comes from years of internalized rules and pattern recognition. AI is increasingly capable of applying those rules, which pushes human value toward context, relationships, verification, and advice. ACCOUNTING MOVES FORWARD The profession has historically reported what already happened. AI makes it possible for accountants to move closer to real-time financial insight and become a co-pilot for what the business should do next. CHAPTERS 00:00 Meet Peter McCarroll 02:41 Train on the Work, Not the Tools 04:52 AI Workflows at Fuel Accountants 07:31 What Should Never Be Automated 10:45 When AI Fails in Accounting 13:02 Making AI More Deterministic 14:28 Choosing Models and Managing Cost 19:02 Client Data, Privacy, and AI 23:46 Accountability for AI-Generated Work 26:51 Is AI Taking Accounting Jobs? 30:10 Filtering AI Noise for Accountants 32:19 Professional Judgment Is Changing 38:45 Hallucinations and Context Engineering 42:31 What Accounting Firms Should Do Now 45:42 Accounting Five Years From Now

  7. Aug 19

    Why China is winning the AI race (Qwen3.8:27B is amazing)

    READ THE FULL EPISODE PAGE https://devmesh.tech/podcast/why-china-is-winning-the-ai-race A 27 billion parameter model should not be competing with frontier AI. But Qwen3.8:27B is making that comparison a lot less ridiculous. In Episode 22 of System Prompt, Peter and Val look at what a model this size can actually do in practical use. Instead of just discussing benchmarks, Peter runs Qwen3.8:27B locally inside Pi Code and gives it a real task during the episode: build a comparative analysis workflow, create test data, work through failures, validate the results, and produce a usable report. It finishes before the episode ends. The bigger question is not whether Qwen replaces frontier models. It is how much work no longer needs a frontier model at all. WHAT WE DISCUSS • Why Qwen3.8:27B matters • Running capable AI locally • Coding and long-running tasks • Tool use and agent workflows • Using local AI for business work • Where smaller models still fall short • Executor models vs heavy reasoning models • Routing harder work to frontier AI • Dense models vs mixture-of-experts • How local AI changes cost and infrastructure KEY TAKEAWAYS 27B MODELS CAN DO REAL WORK Qwen3.8:27B is small enough to run on prosumer hardware while still being capable of coding, tool use, structured analysis, and longer-running tasks. THE HARNESS MATTERS The model does not work alone. Inside Pi Code, Qwen can inspect its environment, create tools, write code, run tests, find problems, and continue working toward a finished result. BUSINESS WORK IS A REAL USE CASE During the episode, Qwen builds a comparative analysis capability from scratch. It creates test data, cleans and normalizes information, performs the analysis, and generates graphs from the results. The output still needs human review, but the model can take meaningful execution work off someone's plate. LOCAL DOES NOT HAVE TO REPLACE FRONTIER The goal is not to eliminate Claude, ChatGPT, or other frontier models. A local model can handle well-defined execution while more ambiguous or difficult work routes to a frontier model when necessary. GOOD SPECS MATTER Qwen performs best when the task is clear. A human or stronger model can define the plan and requirements, then hand execution to the smaller model. That makes routing and task design increasingly important. THE FUTURE IS HYBRID Local models will not win every task, and frontier models are not going away. But as smaller models improve, more work can happen locally while frontier models become the escalation path instead of the default. CHAPTERS 00:00 Episode 22 01:12 Why Qwen3.8:27B? 03:35 Comparing 27B to Frontier AI 05:37 What Can You Actually Do With It? 07:50 Building a Workflow Live 14:09 What Smaller Models Mean 17:11 Local AI Economics 20:26 Internal Business Assistants 23:02 Routing to Frontier Models 26:21 Where Qwen Falls Short 28:36 Do You Need the Best Model? 38:56 Why the Future Is Hybrid 41:55 The Finished Analysis 43:44 Dense vs Mixture-of-Experts 46:36 What Local AI Can Replace 51:01 What 27B Enables Today

  8. Aug 13

    Evals: How Do You Know Which AI Model to Trust?

    READ THE FULL EPISODE PAGE https://devmesh.tech/podcast/how-to-know-which-ai-model-to-trust The AI model at the top of a leaderboard may not be the best model for your system. Because the leaderboard is not testing your system. In Episode 21 of System Prompt, Peter and Val break down AI evals: what benchmarks measure, why the harness matters, and how to test models against the work you actually expect them to do. Peter walks through a custom eval across more than 20 local and open models covering tool calling, extraction, instruction following, and real-world coding tasks. The results were surprising. Smaller models matched or beat much larger ones. Turning reasoning on sometimes made performance worse. The bigger lesson: an eval measures more than the model. Quantization, runtime, token budgets, reasoning settings, parsers, and timeouts can all affect the result. WHAT WE DISCUSS • What AI evals actually measure • Why leaderboards only tell part of the story • How the harness changes model performance • Quantization, runtimes, and configuration • Building evals around real workloads • Tool calling, extraction, instruction following, and coding • Why repetition and consistency matter • Thinking vs non-thinking configurations • Routing tasks to different models • Finding problems in your own system KEY TAKEAWAYS THE BEST MODEL DEPENDS ON THE JOB A benchmark measures performance on a particular test. It does not automatically tell you which model is best for your application. A coding agent, extraction pipeline, chatbot, and tool-using agent all need different things. Start with the workload, then choose the eval. THE HARNESS IS PART OF THE RESULT Models do not operate alone. The harness creates prompts, exposes tools, manages token limits, parses responses, and decides whether a task succeeded. Change the harness, configuration, quantization, or runtime and you can change the result. TEST THE MODEL YOU ARE ACTUALLY RUNNING A full-precision benchmark is useful reference data, but it is not the same experiment as running a Q4 model through a local runtime. Your production configuration is part of the evaluation. REPETITION MATTERS One successful run does not prove reliability. Running tasks multiple times exposes models that score well once but behave inconsistently. For production systems, stability matters. REASONING IS NOT ALWAYS BETTER Thinking modes helped some models and hurt others. In some cases reasoning increased token use, hit time or output budgets, or reduced consistency. The right configuration has to be measured against the task. EVALS ENABLE ROUTING The best architecture may not use one model for everything. A smaller model may handle chat, extraction, or tool calling while another handles coding or harder reasoning. Once you know where each model succeeds and fails, routing stops being guesswork. EVALS TEST YOUR SYSTEM TOO The eval process also exposed problems in Peter's own gateway and harness. Some apparent model failures were really token limits, timeouts, parsing issues, or infrastructure problems. CHAPTERS 00:00 Episode 21 and the 1% 00:48 What Are AI Evals? 02:48 Model Capability and Benchmarks 05:07 Why the Harness Matters 09:48 Quantization and Fair Comparisons 11:30 Building a Custom Eval Suite 19:04 Repetition and Reliability 20:50 The Model Results 23:17 When Thinking Hurts Performance 29:55 Accuracy Versus Token Cost 30:55 Routing Tasks to Different Models 37:00 Evals Finding Bugs in the System 40:36 Closing Thoughts

About

System Prompt is a podcast about what’s actually happening in AI. Not hype. Not surface-level takes. We break down how AI is changing software, SaaS, infrastructure, and the way systems are built focusing on real-world tradeoffs, architecture decisions, and where the value is actually shifting. If you’re building, deploying, or thinking seriously about AI, this is for you.