{"podcast":{"title":"The a16z Show","slug":"the-a16z-show-436525","podcast_index_feed_id":436525,"rss_url":"https://feeds.simplecast.com/JGE3yC0V","website_url":"https://a16z.simplecast.com","image_url":"https://image.simplecastcdn.com/images/0d97354a-306b-45f5-bf26-a8d81eef47ec/ed2664df-9371-438e-8baf-dd2ee0fdde87/3000x3000/thea16zshow-podcastcoverart-3000x3000.jpg?aid=rss_feed","author":"a16z","episode_count":1000,"summary":"The a16z Show discusses tech and culture trends, news, and the future – especially as ‘software eats the world’. It features industry experts, business leaders, and other interesting thinkers and voices from around the world. This show is produced by Andreessen Horowitz (aka “a16z”), a Silicon Valley-based venture capital firm. Multiple episodes are released every week; visit a16z.com for more details and to sign up for our newsletters and other content as well!","last_synced_at":"2026-09-19T16:19:21.546571+00:00","page_url":"https://stenobird.com/podcast/the-a16z-show-436525"},"episode":{"title":"Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan","slug":"who-grades-the-ai-models-ben-horowitz-rayan-krishnan","published_at":"2026-09-09T10:00:00+00:00","page_url":"https://stenobird.com/podcast/the-a16z-show-436525/who-grades-the-ai-models-ben-horowitz-rayan-krishnan","show_page_url":"https://stenobird.com/podcast/the-a16z-show-436525","url":"https://a16z.simplecast.com/episodes/who-grades-the-ai-models-ben-horowitz-rayan-krishnan-uPW5nl18","audio_url":"https://mgln.ai/e/1344/afp-848985-injected.calisto.simplecastaudio.com/3f86df7b-51c6-4101-88a2-550dba782de8/episodes/ba494c31-cd47-4bae-a5d5-89304f960936/audio/128/default.mp3?aid=rss_feed&awCollectionId=3f86df7b-51c6-4101-88a2-550dba782de8&awEpisodeId=ba494c31-cd47-4bae-a5d5-89304f960936&feed=JGE3yC0V","summary":"a16z’s Erik Torenberg, Ben Horowitz, and Jennifer Li sit down with Vals founder and CEO Rayan Krishnan to discuss one of AI’s increasingly difficult problems: how do you actually measure whether a model is getting better? As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. They unpack why self-reported model scores can be misleading, how VALS evaluates models in the hours before a release, and why measuring increasingly agentic systems means testing work that can unfold over hours, days, or even weeks. They also explore why evals are becoming critical for enterprises trying to understand the ROI of AI, what happens if token spend begins to rival employee salaries, and how evaluations could eventually provide a shared language for everything from model routing and recursive self-improvement to AI policy and international coordination.","meta_description":"a16z’s Erik Torenberg, Ben Horowitz, and Jennifer Li sit down with Vals founder and CEO Rayan Krishnan to discuss one of AI’s increasingly difficult probl…","key_points":[],"chapters":[],"topics":[],"duration_seconds":2385,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/the-a16z-show-436525/episodes/who-grades-the-ai-models-ben-horowitz-rayan-krishnan/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/the-a16z-show-436525/who-grades-the-ai-models-ben-horowitz-rayan-krishnan.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}