{"podcast":{"title":"In The Blink of AI with Georgie Healy","slug":"in-the-blink-of-ai-with-georgie-healy-7077303","podcast_index_feed_id":7077303,"rss_url":"https://feeds.captivate.fm/blink-of-ai-dayone/","website_url":"https://dayone.fm/show/in-the-blink-of-ai","image_url":"https://artwork.captivate.fm/10e2b33a-c7c8-4fc9-ac4b-176c00c2e39f/powered-by.png","author":"Day One®","episode_count":81,"summary":"Stay ahead of the curve in the rapidly changing world of technology with the In the Blink of AI podcast Host of the show Georgie Healy leverages 15 years in tech with a vibrant Aussie sense of humor to interview leading experts in Artificial Intelligence to unpack AI’s transformative potential and provide weekly commentary on the latest headlines. Tune in for candid conversations as the rapid speed of technology navigates innovation and ethics. The podcast's mission is to demystify the AI jargon and provide real human insights for everyone from the AI-curious to seasoned experts. Hosted by Georgie Healy, In the Blink of AI is a Day One® show. Day One is the podcast network dedicated to founders, operators, and investors. Follow In the Blink of AI through Day One on LinkedIn Sign up to get your weekly insights into the up-and-coming AI startups.","last_synced_at":"2026-07-17T08:19:48.857875+00:00","page_url":"https://stenobird.com/podcast/in-the-blink-of-ai-with-georgie-healy-7077303"},"episode":{"title":"How AI Models Are Really Judged, with Peter Gostev (Arena / LMArena)","slug":"how-ai-models-are-really-judged-with-peter-gostev-arena-lmarena","published_at":"2026-07-02T20:00:00+00:00","page_url":"https://stenobird.com/podcast/in-the-blink-of-ai-with-georgie-healy-7077303/how-ai-models-are-really-judged-with-peter-gostev-arena-lmarena","show_page_url":"https://stenobird.com/podcast/in-the-blink-of-ai-with-georgie-healy-7077303","url":"https://dayone.fm/show/in-the-blink-of-ai","audio_url":"https://dts.podtrac.com/redirect.mp3/prfx.byspotify.com/e/episodes.captivate.fm/episode/b2e0cdfb-7692-4494-b448-fde308c39f98.mp3","summary":"Peter Gostev is head of AI capabilities at Arena (LMArena), the community-based platform where millions of real people vote in blind tests to rank AI models, born out of research at UC Berkeley. Before Arena, Peter was Head of AI at Moonpig and built a large following sharing hands-on explorations of what the latest models can actually do. He joins Georgie Healy from London for a genuinely nerdy, insider look at how models are judged and where the frontier is heading. In this episode, Peter explains the difference between static benchmarks and human judgment, and why a model can pass every test you write and still produce something that looks completely awful. He breaks down the current state of the leaderboards, why Anthropic's models are dominating and how that tracks with real world adoption, and gives a sharp comparison of the top Western models, including why Anthropic's non-reasoning models are exceptional while OpenAI's strength lies in deep reasoning. Georgie and Peter get into why people aren't using Chinese models more despite their quality, the economics behind AI pricing and how enterprise usage is priced very differently from consumer subscriptions, why release cadence matters as much as capability, and what the wave of data centre investment means for the models arriving next. Along the way there's a fond detour on Opus 3 as the model you could talk to for hours, and why better models can sometimes feel worse. Tune in for a clear-eyed, hype-free guide to how AI models are really evaluated, straight from someone who watches the charts move in real time. Mentioned in this episode: Deel x PX_Post Intro","meta_description":"Peter Gostev is head of AI capabilities at Arena (LMArena), the community-based platform where millions of real people vote in blind tests to rank AI mode…","key_points":[],"chapters":[],"topics":[],"duration_seconds":3483,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/in-the-blink-of-ai-with-georgie-healy-7077303/episodes/how-ai-models-are-really-judged-with-peter-gostev-arena-lmarena/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/in-the-blink-of-ai-with-georgie-healy-7077303/how-ai-models-are-really-judged-with-peter-gostev-arena-lmarena.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}