{"podcast":{"title":"Chain of Thought | AI Agents, Infrastructure & Engineering","slug":"chain-of-thought-ai-agents","podcast_index_feed_id":7074333,"rss_url":"https://feeds.transistor.fm/chain-of-thought","website_url":"https://www.chainofthought.show/","image_url":"https://img.transistorcdn.com/L9UtZvVIeHCUeyWb9gm7IMgzws11SpR5smiF7YG89n4/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS85ZWU3/Mjk4NDcwNGRiOGI4/NWUwMTQ0ZjA1NjM3/ZWU0Ny5qcGc.jpg","author":"Conor Bronsdon","episode_count":73,"summary":"AI is reshaping infrastructure, strategy, and entire industries. Chain of Thought is the podcast where builders reason through what's changing. Host Conor Bronsdon sits down with the engineers and founders shipping AI in production to get past the hype into what's working and what isn't. Episodes cover model infrastructure, inference, agent frameworks, evaluation, and developer tools. Guests have come from NVIDIA, Google DeepMind, AMD, Databricks, Vercel, and more. Every episode carries a full transcript and show notes at chainofthought.show. New episodes weekly. Conor Bronsdon is an independent consultant and angel investor in AI infrastructure and developer tools. He led technical ecosystem at Modular, acquired by Qualcomm in 2026; led developer awareness at Galileo, acquired by Cisco; and ran developer marketing at LinearB, where he was GM of the Dev Interrupted podcast and community. Views expressed by the host and guests are their own.","last_synced_at":"2026-09-10T00:20:07.137048+00:00","page_url":"https://stenobird.com/podcast/chain-of-thought-ai-agents"},"episode":{"title":"Explaining Eval Engineering | Galileo's Vikram Chatterji","slug":"explaining-eval-engineering-galileo-s-vikram-chatterji","published_at":"2025-12-19T10:00:00+00:00","page_url":"https://stenobird.com/podcast/chain-of-thought-ai-agents/explaining-eval-engineering-galileo-s-vikram-chatterji","show_page_url":"https://stenobird.com/podcast/chain-of-thought-ai-agents","url":"https://share.transistor.fm/s/28aaae24","audio_url":"https://media.transistor.fm/28aaae24/b3da0aa0.mp3","summary":"You've heard of evaluations—but eval engineering is the difference between AI that ships and AI that's stuck in prototype.Most teams still treat evals like unit tests: write them once, check a box, move on. But when you're deploying agents that make real decisions, touch real customers, and cost real money, those one-time tests don't cut it. The companies actually shipping production AI at scale have figured out something different—they've turned evaluations into infrastructure, into IP, into the layer where domain expertise becomes executable governance.Vikram Chatterji, CEO and Co-founder of Galileo, returns to Chain of Thought to break down eval engineering: what it is, why it's becoming a dedicated discipline, and what it takes to actually make it work. Vikram shares why generic evals are plateauing, how continuous learning loops drive accuracy, and why he predicts \"eval engineer\" will become as common a role as \"prompt engineer\" once was.In this conversation, Conor and Vikram explore:Why treating evals as infrastructure—not checkboxes—separates production AI from prototypesThe plateau problem: why generic LLM-as-a-judge metrics can't break 90% accuracyHow continuous human feedback loops improve eval precision over timeThe emerging \"eval engineer\" role and what the job actually looks likeWhy 60-70% of AI engineers' time is already spent on evalsWhat multi-agent systems mean for the future of evaluationVikram's framework for baking trust AND control into agentic applicationsPlus: Conor shares news about his move to Modular and what it means for Chain of Thought going forward.Chapters:00:00 – Introduction: Why Evals Are Becoming IP01:37 – What Is Eval Engineering?04:24 – The Eval Engineering Course for Developers05:24 – Generic Evals Are Plateauing08:21 – Continuous…","meta_description":"You've heard of evaluations—but eval engineering is the difference between AI that ships and AI that's stuck in prototype.Most teams still treat evals lik…","key_points":[],"chapters":[],"topics":[],"duration_seconds":2235,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/chain-of-thought-ai-agents/episodes/explaining-eval-engineering-galileo-s-vikram-chatterji/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/chain-of-thought-ai-agents/explaining-eval-engineering-galileo-s-vikram-chatterji.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}