Episode

Cheaper Tokens, Bigger Bills

Podcast
YPO Technology Network AI Brief
Published
Jun 22, 2026
Duration seconds
506
Processing state
not_requested
Canonical source
https://rss.com/podcasts/ypo-technology-network-ai-brief/2932322
Audio
https://content.rss.com/episodes/382927/2932322/ypo-technology-network-ai-brief/2026_06_21_16_21_38_fdf154cb-5285-4a92-98ed-3eeb827fff47.mp3
JSON
/v1/public/podcasts/ypo-technology-network-ai-brief-7728971/episodes/cheaper-tokens-bigger-bills
Markdown
/podcast/ypo-technology-network-ai-brief-7728971/cheaper-tokens-bigger-bills.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/ypo-technology-network-ai-brief-7728971/episodes/cheaper-tokens-bigger-bills/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/ypo-technology-network-ai-brief-7728971/cheaper-tokens-bigger-bills.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

The strategic AI question is no longer "which model do we use." It's "where does the model run, and who pays for the tokens." This week the AI inference startup Baseten raised roughly $1.5 billion at up to a $13 billion valuation for the unglamorous business of running other companies' models. Meanwhile token prices are collapsing about 10x a year, and enterprise AI bills are going up anyway. In this episode, Stephen Forte unpacks the inference economy and what it means for your business: The inference gold rush — why investors value the company that runs models more than many that build them, and why inference is 80-90% of a model's lifetime cost. The land grab — Amazon selling its Trainium chips to challenge Nvidia, and Alphabet's $84.75 billion raise to fund AI capex. The pricing paradox — "LLMflation" makes tokens ~10x cheaper a year, yet the Jevons paradox and the new "thinking tax" of reasoning models send total bills higher. The counter-move — open-weight models running locally on your own hardware, Apple's new "zero token cost" Core AI, and how to think about cloud vs. local as a cost-structure decision. Two concrete moves for the quarter — build multi-model routing, and budget for usage growth, not the falling unit price. Sources: Baseten ~$1.5B round at up to $13B — TechCrunch Amazon to sell Trainium chips externally — TechCrunch Alphabet $84.75B equity offering for AI — Intellectia LLMflation and falling inference costs — a16z Jevons paradox and rising enterprise AI spend — GUUTs / FinOps Apple Core AI at WWDC26 — Let's Data Science The AI Brief from the YPO Technology Network is a daily executive briefing on the AI developments that matter to business leaders. Hosted by Stephen Forte.