Episode

Small Models, Massive Wins: The New Shopify AI Formula

Podcast
Beyond The Pilot: Enterprise AI in Action
Published
Jun 24, 2026
Duration seconds
2650
Processing state
not_requested
Canonical source
https://traffic.megaphone.fm/UTEAU7770800394.mp3
Audio
https://traffic.megaphone.fm/UTEAU7770800394.mp3
JSON
/v1/public/podcasts/beyond-the-pilot-enterprise-ai-in-action-7531113/episodes/small-models-massive-wins-the-new-shopify-ai-formula
Markdown
/podcast/beyond-the-pilot-enterprise-ai-in-action-7531113/small-models-massive-wins-the-new-shopify-ai-formula.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/beyond-the-pilot-enterprise-ai-in-action-7531113/episodes/small-models-massive-wins-the-new-shopify-ai-formula/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/beyond-the-pilot-enterprise-ai-in-action-7531113/small-models-massive-wins-the-new-shopify-ai-formula.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

Shopify's distillation pipeline cuts production AI costs by up to 30x — and in some cases, the smaller model outperforms the frontier model on the narrow task. That's not a trade-off. That's a win on accuracy, latency, and cost simultaneously. Farhan Thawar, VP & Head of Engineering at Shopify, runs AI across one of the largest commerce platforms on earth. In this episode, he breaks down the exact infrastructure decisions Shopify made to avoid being locked into any single model provider — and why 29% of enterprise AI projects die from token costs, not model failure. Shopify built an internal LLM proxy that routes tokens across every major provider, enabling automatic failover when any one goes down. On top of that, their Universal Distillation Platform (UDP) lets any R&D team distill a frontier model (Opus 4, GPT-5+) down to a fine-tuned open source model (Qwen and others) for a specific subtask — in roughly a day, with evals baked in. Results range from 2x to 30x cheaper, faster, and more accurate than calling the frontier API for everything. Shopify currently runs roughly half a dozen of these distilled models in production, with more being added. Farhan also details River, Shopify's internal agentic substrate — a public-only Slack agent that queries their data warehouse, reads their PM system, and improves its own answers when engineers jump in to correct it. Plus: how Shopify governs AI-generated code at scale, why they killed their token leaderboard, how UCP is positioning Shopify's catalog for agentic commerce, and what a two-to-three year horizon looks like when agents start holding spending budgets and buying autonomously. 🎙️ GUEST: Farhan Thawar | VP & Head of Engineering, Shopify 🎙️ HOST: Sam Witteveen | VentureBeat __ If you enjoy these conversations, you ne…