Episode

How AI Models Are Really Judged, with Peter Gostev (Arena / LMArena)

Podcast
In The Blink of AI with Georgie Healy
Published
Jul 2, 2026
Duration seconds
3483
Processing state
not_requested
Canonical source
https://dayone.fm/show/in-the-blink-of-ai
Audio
https://dts.podtrac.com/redirect.mp3/prfx.byspotify.com/e/episodes.captivate.fm/episode/b2e0cdfb-7692-4494-b448-fde308c39f98.mp3
JSON
/v1/public/podcasts/in-the-blink-of-ai-with-georgie-healy-7077303/episodes/how-ai-models-are-really-judged-with-peter-gostev-arena-lmarena
Markdown
/podcast/in-the-blink-of-ai-with-georgie-healy-7077303/how-ai-models-are-really-judged-with-peter-gostev-arena-lmarena.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/in-the-blink-of-ai-with-georgie-healy-7077303/episodes/how-ai-models-are-really-judged-with-peter-gostev-arena-lmarena/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/in-the-blink-of-ai-with-georgie-healy-7077303/how-ai-models-are-really-judged-with-peter-gostev-arena-lmarena.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

Peter Gostev is head of AI capabilities at Arena (LMArena), the community-based platform where millions of real people vote in blind tests to rank AI models, born out of research at UC Berkeley. Before Arena, Peter was Head of AI at Moonpig and built a large following sharing hands-on explorations of what the latest models can actually do. He joins Georgie Healy from London for a genuinely nerdy, insider look at how models are judged and where the frontier is heading. In this episode, Peter explains the difference between static benchmarks and human judgment, and why a model can pass every test you write and still produce something that looks completely awful. He breaks down the current state of the leaderboards, why Anthropic's models are dominating and how that tracks with real world adoption, and gives a sharp comparison of the top Western models, including why Anthropic's non-reasoning models are exceptional while OpenAI's strength lies in deep reasoning. Georgie and Peter get into why people aren't using Chinese models more despite their quality, the economics behind AI pricing and how enterprise usage is priced very differently from consumer subscriptions, why release cadence matters as much as capability, and what the wave of data centre investment means for the models arriving next. Along the way there's a fond detour on Opus 3 as the model you could talk to for hours, and why better models can sometimes feel worse. Tune in for a clear-eyed, hype-free guide to how AI models are really evaluated, straight from someone who watches the charts move in real time. Mentioned in this episode: Deel x PX_Post Intro