Episode

Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan

Podcast
The a16z Show
Published
Sep 9, 2026
Duration seconds
2385
Processing state
not_requested
Canonical source
https://a16z.simplecast.com/episodes/who-grades-the-ai-models-ben-horowitz-rayan-krishnan-uPW5nl18
Audio
https://mgln.ai/e/1344/afp-848985-injected.calisto.simplecastaudio.com/3f86df7b-51c6-4101-88a2-550dba782de8/episodes/ba494c31-cd47-4bae-a5d5-89304f960936/audio/128/default.mp3?aid=rss_feed&awCollectionId=3f86df7b-51c6-4101-88a2-550dba782de8&awEpisodeId=ba494c31-cd47-4bae-a5d5-89304f960936&feed=JGE3yC0V
JSON
/v1/public/podcasts/the-a16z-show-436525/episodes/who-grades-the-ai-models-ben-horowitz-rayan-krishnan
Markdown
/podcast/the-a16z-show-436525/who-grades-the-ai-models-ben-horowitz-rayan-krishnan.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/the-a16z-show-436525/episodes/who-grades-the-ai-models-ben-horowitz-rayan-krishnan/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/the-a16z-show-436525/who-grades-the-ai-models-ben-horowitz-rayan-krishnan.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

a16z’s Erik Torenberg, Ben Horowitz, and Jennifer Li sit down with Vals founder and CEO Rayan Krishnan to discuss one of AI’s increasingly difficult problems: how do you actually measure whether a model is getting better? As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. They unpack why self-reported model scores can be misleading, how VALS evaluates models in the hours before a release, and why measuring increasingly agentic systems means testing work that can unfold over hours, days, or even weeks. They also explore why evals are becoming critical for enterprises trying to understand the ROI of AI, what happens if token spend begins to rival employee salaries, and how evaluations could eventually provide a shared language for everything from model routing and recursive self-improvement to AI policy and international coordination.