# Haize Labs with Leonard Tang - Weaviate Podcast #121! Page: https://stenobird.com/podcast/weaviate-podcast-6288219/haize-labs-with-leonard-tang-weaviate-podcast-121 Text version: https://stenobird.com/podcast/weaviate-podcast-6288219/haize-labs-with-leonard-tang-weaviate-podcast-121.md Podcast: [Weaviate Podcast](https://stenobird.com/podcast/weaviate-podcast-6288219) Published: 2025-05-12T14:56:52+00:00 Episode link: https://podcasters.spotify.com/pod/show/weaviate/episodes/Haize-Labs-with-Leonard-Tang---Weaviate-Podcast-121-e32mts3 Audio file: https://anchor.fm/s/cffc3468/podcast/play/102511939/https%3A%2F%2Fd3ctxlq1ktw2nl.cloudfront.net%2Fstaging%2F2025-4-11%2F9301155a-7b15-111f-e54c-5c1d7c0d9caf.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/weaviate-podcast-6288219/episodes/haize-labs-with-leonard-tang-weaviate-podcast-121 Duration seconds: 3255 ## Resource How do you ensure your AI systems actually do what you expect them to do? Leonard Tang takes us deep into the revolutionary world of AI evaluation with concrete techniques you can apply today. Learn how Haize Labs is transforming AI testing through "scaling judge-time compute" - stacking weaker models to effectively evaluate stronger ones. Leonard unpacks the game-changing Verdict library that outperforms frontier models by 10-20% while dramatically reducing costs. Discover practical insights on creating contrastive evaluation sets that extract maximum signal from human feedback, implementing debate-based judging systems, and building custom reward models that align with enterprise needs. The conversation reveals powerful nuggets like using randomized agent debates to achieve consensus and lightweight guardrail models that run alongside inference. Whether you're developing AI applications or simply fascinated by how we'll ensure increasingly powerful AI systems perform as expected, this episode delivers immediate value with techniques you can implement right away, philosophical perspectives on AI safety, and a glimpse into the future of evaluation that will fundamentally shape how AI evolves. ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/weaviate-podcast-6288219/episodes/haize-labs-with-leonard-tang-weaviate-podcast-121/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/weaviate-podcast-6288219/haize-labs-with-leonard-tang-weaviate-podcast-121.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.