In this episode, Sheamus McGovern speaks with Parth Sareen, a leading software engineer at Ollama, about local AI, open models, and the agent harness.
Parth works across inference, sampling, structured outputs, APIs, and agent tooling at Ollama. He created Ollama Launch, led the OpenClaw launch, and has worked on core systems that make Ollama easier for developers to use.
The conversation starts with what Ollama does and why it has become a go-to tool for running open models locally and in the cloud. It then moves into the bigger technical theme of the episode: the agent harness, or the tools, memory, loops, file access, web access, and execution environment around the model.
The episode also covers local-first AI, Ollama Launch, structured outputs, tool design, verification loops, context engineering, and how coding agents are changing software engineering workflows.
Key Topics Covered:
What Ollama is and why developers use it to run open models locally and in the cloud
Ollama’s move from running models to connecting them to real applications
How Ollama Launch reduces setup friction for agentic tools and coding assistants
What an agent harness is and why it matters for building useful AI agents
Local models for privacy, offline workflows, and smaller development tasks
When cloud models are still better for heavier coding-agent workloads
Why tool design is critical for agent reliability
How structured outputs help local models produce consistent results
Verification loops, scripts, tests, and benchmarks for more reliable agents
How coding agents are changing the role of the software engineer
Where computer use agents may be headed next
Memorable Outtakes:
“We don’t just run the models anymore, we also try and connect them to people’s actual applications where they’re trying to make use of these models to do real work.”
“A harness is really everything that kind of holds an agent.”
References & Resources:
Speaker:
Parth Sareen: https://www.linkedin.com/in/parthsareen/
Parth Sareen personal site and blog: https://parthsareen.com
Ollama:
Ollama: https://ollama.com/
Ollama GitHub: https://github.com/ollama/ollama
Ollama Docs: https://ollama.com/docs
Articles and technical references mentioned:
Parth Sareen writings: https://parthsareen.com/writings/
Yet Another Generalist: https://parthsareen.com/writings/era-of-generalists/
Building reliable AI agents: https://parthsareen.com/writings/building-reliable-ai-agents/
Sampling and structured outputs in LLMs: https://parthsareen.com/writings/sampling/
OpenClaw: https://ollama.com/blog/openclaw
Terminal-Bench: https://www.tbench.ai/
Amp Code article, “How to Build an Agent”: https://ampcode.com/notes/how-to-build-an-agent
Other Resources
OpenClaw: https://docs.openclaw.ai/
Hermes Agent: https://hermes-agent.nousresearch.com
Pi Coding Agent Harness: https://pi.dev
Claude Code: https://claude.com/product/claude-code
OpenAI Codex CLI: https://developers.openai.com/codex/cli
OpenCode: https://opencode.ai/
MLX: https://github.com/ml-explore/mlx
Book recommendation: The Alchemist by Paulo Coelho: https://www.paulocoelho.com/books/
Sponsored by:
This episode was sponsored by:
ODSC AI West 2026 – The Leading AI Builders Conference
Join us in San Francisco from October 27th–29th for expert-led sessions on Agentic AI, AI Engineering, Data Science, Machine Learning, LLMOps, and AI-driven automation.
Learn more: https://odsc.ai/west
ODSC AI Engineering Accelerator – From Using AI to Building It
Seven weeks of live, cohort-based training in LLMs, AI assistants, agents, and AI engineering. Ship a working AI agent or assistant by the end. Summer and Fall cohorts available. Group pricing for teams of 3+.
Learn more: https://odsc.ai/west/accelerator
Information
- Show
- FrequencyUpdated Weekly
- PublishedMay 27, 2026 at 4:00 AM UTC
- Season1
- Episode115
