Episode

[REDACTED] Episode 4: We Stopped Using Claude Code Mid-Build. Here's What We Built Instead

Podcast
NC Tweener Talks
Published
Jun 17, 2026
Duration seconds
2496
Processing state
not_requested
Canonical source
https://share.transistor.fm/s/6581314c
Audio
https://media.transistor.fm/6581314c/678637a5.mp3
JSON
/v1/public/podcasts/nc-tweener-talks-7053588/episodes/redacted-episode-4-we-stopped-using-claude-code-mid-build-here-s-what-we-built-instead
Markdown
/podcast/nc-tweener-talks-7053588/redacted-episode-4-we-stopped-using-claude-code-mid-build-here-s-what-we-built-instead.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/nc-tweener-talks-7053588/episodes/redacted-episode-4-we-stopped-using-claude-code-mid-build-here-s-what-we-built-instead/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/nc-tweener-talks-7053588/redacted-episode-4-we-stopped-using-claude-code-mid-build-here-s-what-we-built-instead.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

Redacted is the show that doesn't clean things up before hitting record. Episode 4 is a double build session: Taylor Cotner walks through the multi-agent HubSpot cleanup pipeline he's been iterating on for weeks, now running on the Anthropic SDK with Claude Code out of the loop, and David Shaner demos how he used Claude Design and Claude Code to rebuild Offline's partner landing page from scratch. Most of the episode is screen sharing, so pull it up on YouTube. What We Cover No Claude Code in the loop: Taylor stopped using Claude Code as an agent orchestrator in his HubSpot pipeline, not as his coding tool (he’s still building the app with Claude Code), but as a decision-maker in the middle of a workflow. Removing it gave him full control over inputs and outputs at every step. Custom eval system built from scratch: Taylor built an eval page that looks like an Excel grid, models as columns, test cases as rows , to measure Haiku, Sonnet, GPT-5, and GPT-5 Mini against real messy HubSpot data. Each cell shows pass, fail, and cost. GPT-5 Mini at 10–20× less cost: For the lead qualifier agent, Sonnet evals cost $1.00 per run. GPT-5 Mini costs $0.05. “I can live with that. 10X less the cost.” For the core cleanup evals: $1.50 for Sonnet versus $0.14 on GPT-5 Mini. $20 for 133 million tokens overnight: Using the Vercel AI Gateway — which lets you swap any model without changing your code, Taylor ran 200 HubSpot restaurant cleanups in a single night for $20 total. Self-grading pipeline: The pipeline grades its own output after every cleanup run. If a job comes back below an A, it automatically spawns a new run with Sonnet, no human catch required. A B grade on 101 Craft Kitchen auto-escalated and came back with an A. Real mess-ups make the best evals: Almost every eval case cam…