Episode

Sonnet 5 review: I ran 64 generations to find out if it's worth it

Podcast
How I AI
Published
Jun 30, 2026
Duration seconds
1556
Processing state
not_requested
Canonical source
https://podcasters.spotify.com/pod/show/pen-name/episodes/Sonnet-5-review-I-ran-64-generations-to-find-out-if-its-worth-it-e3lga05
Audio
https://anchor.fm/s/1035b1568/podcast/play/122217925/https%3A%2F%2Fd3ctxlq1ktw2nl.cloudfront.net%2Fstaging%2F2026-5-30%2F427131012-44100-2-00e3dda044fe6.mp3
JSON
/v1/public/podcasts/how-i-ai-7304222/episodes/sonnet-5-review-i-ran-64-generations-to-find-out-if-it-s-worth-it
Markdown
/podcast/how-i-ai-7304222/sonnet-5-review-i-ran-64-generations-to-find-out-if-it-s-worth-it.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/how-i-ai-7304222/episodes/sonnet-5-review-i-ran-64-generations-to-find-out-if-it-s-worth-it/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/how-i-ai-7304222/sonnet-5-review-i-ran-64-generations-to-find-out-if-it-s-worth-it.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

I’ve been testing every major frontier model release since the start of the year, and when Anthropic dropped Sonnet 5, I wanted more than a vibe check. I got tired of one-off tests I couldn’t repeat or compare over time, so I built something better: the How I AI Bench, a repeatable eval harness I constructed live using Claude Code while recording this episode. I ran Sonnet 5 blind against four other frontier models (Sonnet 4.6, Opus 4.8, GPT-5.5, and Gemini 3 Pro) across PRD quality, prototype generation, agentic task completion, and agent personality. The results were not what I expected. What you’ll learn: What Anthropic claims Sonnet 5 improves over Sonnet 4.6, and where the benchmark data actually backs that up How I built the How I AI Bench in under 45 minutes using Claude Code, starting from my own stored session history Why I combined human vibe scoring (70%) with LLM as judge scoring (30%) instead of trusting either alone How to set up a local HTML scoring page so you can rate AI outputs on gut feel and export those scores as JSON Which model I recommend for PRDs, which for complex prototypes, and which for chatting with an agent daily — Brought to you by: Runway —The creative AI platform for images, video and more Hyperagent —Deploy fleets of agents that handle real work — In this episode, we cover: (00:00) Sonnet 5 is out (01:55) What Anthropic claims (04:02) Why I’m done with one-off vibe checks (05:05) Building the How I AI Bench live with Claude Code (07:42) The scoring system (10:43) Agent voice eval (11:57) Quick recap (13:58) Results: The How I AI index leaderboard (21:21) What I’m improving for the next run (22:16) Generating a Claire-weighted index (23:53) Model-by-task recommendations — Tools referenced: • Claude Sonnet 5: https://www.anthropic.com/n…