Episode
Evals are the new PRD. Here is the playbook with the CEO of the leader in the space (Ankur Goyal, Founder and CEO, Braintrust)
- Podcast
- The Growth Podcast
- Published
- Mar 20, 2026
- Duration seconds
- 3117
- Processing state
not_requested- Canonical source
- https://www.news.aakashg.com/p/ankur-goyal-podcast
Actions
POST https://stenobird.com/v1/public/podcasts/the-growth-podcast-6993153/episodes/evals-are-the-new-prd-here-is-the-playbook-with-the-ceo-of-the-leader-in-the-space-ankur-goyal-founder-and-ceo-braintrust/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/the-growth-podcast-6993153/evals-are-the-new-prd-here-is-the-playbook-with-the-ceo-of-the-leader-in-the-space-ankur-goyal-founder-and-ceo-braintrust.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
Today’s episode Most PMs treat evals like a quality gate. Something you run right before shipping, just to check the box. That is backwards . The best AI product teams treat evals as the starting point . They write the eval before the prompt. They iterate on the scoring function before the model. They use failing evals as a roadmap. That shift is what today’s episode is about. I sat down with Ankur Goyal, Founder and CEO of Braintrust. It is the eval platform used by Replit, Vercel, Airtable, Ramp, Zapier, and Notion. Braintrust just announced its Series B at an $800 million valuation . Users are running 10x more evals than this time last year. People log more data per day now than they did in the entire first year the product existed. In this episode, we build an eval entirely from scratch . Live. No pre-written prompts, no pre-written data. We connect to Linear’s MCP server, generate test data, write a scoring function, and iterate until the score goes from 0 to 0.75. Plus, we cover the complete eval playbook for PMs : If you want access to my AI tool stack - Dovetail, Arize, Linear, Descript, Reforge Build, DeepSky, Relay.app, Magic Patterns, Speechify, and Mobbin - grab Aakash’s bundle . If you want my PM Operating System in Claude Code, click here . ---- Check out the conversation on Apple , Spotify , and YouTube . Brought to you by: * Kameleoon : Leading AI experimentation platform * Testkube : Leading test orchestration platform * Pendo : The #1 software experience management platform * Bolt : Ship AI-powered products 10x faster * Product Faculty : Get $550 off their #1 AI PM Certification with my link ---- Key Takeaways: 1. Vibe checks are evals - When you look at an AI output and intuit whether it is good or bad, you are using your brain as a scoring function.…