Episode
Small Models, Massive Wins: The New Shopify AI Formula
- Published
- Jun 24, 2026
- Duration seconds
- 2650
- Processing state
not_requested- Canonical source
- https://traffic.megaphone.fm/UTEAU7770800394.mp3
Actions
POST https://stenobird.com/v1/public/podcasts/beyond-the-pilot-enterprise-ai-in-action-7531113/episodes/small-models-massive-wins-the-new-shopify-ai-formula/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/beyond-the-pilot-enterprise-ai-in-action-7531113/small-models-massive-wins-the-new-shopify-ai-formula.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
Shopify's distillation pipeline cuts production AI costs by up to 30x — and in some cases, the smaller model outperforms the frontier model on the narrow task. That's not a trade-off. That's a win on accuracy, latency, and cost simultaneously. Farhan Thawar, VP & Head of Engineering at Shopify, runs AI across one of the largest commerce platforms on earth. In this episode, he breaks down the exact infrastructure decisions Shopify made to avoid being locked into any single model provider — and why 29% of enterprise AI projects die from token costs, not model failure. Shopify built an internal LLM proxy that routes tokens across every major provider, enabling automatic failover when any one goes down. On top of that, their Universal Distillation Platform (UDP) lets any R&D team distill a frontier model (Opus 4, GPT-5+) down to a fine-tuned open source model (Qwen and others) for a specific subtask — in roughly a day, with evals baked in. Results range from 2x to 30x cheaper, faster, and more accurate than calling the frontier API for everything. Shopify currently runs roughly half a dozen of these distilled models in production, with more being added. Farhan also details River, Shopify's internal agentic substrate — a public-only Slack agent that queries their data warehouse, reads their PM system, and improves its own answers when engineers jump in to correct it. Plus: how Shopify governs AI-generated code at scale, why they killed their token leaderboard, how UCP is positioning Shopify's catalog for agentic commerce, and what a two-to-three year horizon looks like when agents start holding spending budgets and buying autonomously. 🎙️ GUEST: Farhan Thawar | VP & Head of Engineering, Shopify 🎙️ HOST: Sam Witteveen | VentureBeat __ If you enjoy these conversations, you ne…