Episode

Evals are the new PRD. Here is the playbook with the CEO of the leader in the space (Ankur Goyal, Founder and CEO, Braintrust)

Podcast
The Growth Podcast
Published
Mar 20, 2026
Duration seconds
3117
Processing state
not_requested
Canonical source
https://www.news.aakashg.com/p/ankur-goyal-podcast
Audio
https://api.substack.com/feed/podcast/191458506/1ef95026013ab2632eb95c919a9013c5.mp3
JSON
/v1/public/podcasts/the-growth-podcast-6993153/episodes/evals-are-the-new-prd-here-is-the-playbook-with-the-ceo-of-the-leader-in-the-space-ankur-goyal-founder-and-ceo-braintrust
Markdown
/podcast/the-growth-podcast-6993153/evals-are-the-new-prd-here-is-the-playbook-with-the-ceo-of-the-leader-in-the-space-ankur-goyal-founder-and-ceo-braintrust.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/the-growth-podcast-6993153/episodes/evals-are-the-new-prd-here-is-the-playbook-with-the-ceo-of-the-leader-in-the-space-ankur-goyal-founder-and-ceo-braintrust/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/the-growth-podcast-6993153/evals-are-the-new-prd-here-is-the-playbook-with-the-ceo-of-the-leader-in-the-space-ankur-goyal-founder-and-ceo-braintrust.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

Today’s episode Most PMs treat evals like a quality gate. Something you run right before shipping, just to check the box. That is backwards . The best AI product teams treat evals as the starting point . They write the eval before the prompt. They iterate on the scoring function before the model. They use failing evals as a roadmap. That shift is what today’s episode is about. I sat down with Ankur Goyal, Founder and CEO of Braintrust. It is the eval platform used by Replit, Vercel, Airtable, Ramp, Zapier, and Notion. Braintrust just announced its Series B at an $800 million valuation . Users are running 10x more evals than this time last year. People log more data per day now than they did in the entire first year the product existed. In this episode, we build an eval entirely from scratch . Live. No pre-written prompts, no pre-written data. We connect to Linear’s MCP server, generate test data, write a scoring function, and iterate until the score goes from 0 to 0.75. Plus, we cover the complete eval playbook for PMs : If you want access to my AI tool stack - Dovetail, Arize, Linear, Descript, Reforge Build, DeepSky, Relay.app, Magic Patterns, Speechify, and Mobbin - grab Aakash’s bundle . If you want my PM Operating System in Claude Code, click here . ---- Check out the conversation on Apple , Spotify , and YouTube . Brought to you by: * Kameleoon : Leading AI experimentation platform * Testkube : Leading test orchestration platform * Pendo : The #1 software experience management platform * Bolt : Ship AI-powered products 10x faster * Product Faculty : Get $550 off their #1 AI PM Certification with my link ---- Key Takeaways: 1. Vibe checks are evals - When you look at an AI output and intuit whether it is good or bad, you are using your brain as a scoring function.…