# Small Models, Massive Wins: The New Shopify AI Formula Page: https://stenobird.com/podcast/beyond-the-pilot-enterprise-ai-in-action-7531113/small-models-massive-wins-the-new-shopify-ai-formula Text version: https://stenobird.com/podcast/beyond-the-pilot-enterprise-ai-in-action-7531113/small-models-massive-wins-the-new-shopify-ai-formula.md Podcast: [Beyond The Pilot: Enterprise AI in Action](https://stenobird.com/podcast/beyond-the-pilot-enterprise-ai-in-action-7531113) Published: 2026-06-24T10:29:00+00:00 Episode link: https://traffic.megaphone.fm/UTEAU7770800394.mp3 Audio file: https://traffic.megaphone.fm/UTEAU7770800394.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/beyond-the-pilot-enterprise-ai-in-action-7531113/episodes/small-models-massive-wins-the-new-shopify-ai-formula Duration seconds: 2650 ## Resource Shopify's distillation pipeline cuts production AI costs by up to 30x β€” and in some cases, the smaller model outperforms the frontier model on the narrow task. That's not a trade-off. That's a win on accuracy, latency, and cost simultaneously. Farhan Thawar, VP & Head of Engineering at Shopify, runs AI across one of the largest commerce platforms on earth. In this episode, he breaks down the exact infrastructure decisions Shopify made to avoid being locked into any single model provider β€” and why 29% of enterprise AI projects die from token costs, not model failure. Shopify built an internal LLM proxy that routes tokens across every major provider, enabling automatic failover when any one goes down. On top of that, their Universal Distillation Platform (UDP) lets any R&D team distill a frontier model (Opus 4, GPT-5+) down to a fine-tuned open source model (Qwen and others) for a specific subtask β€” in roughly a day, with evals baked in. Results range from 2x to 30x cheaper, faster, and more accurate than calling the frontier API for everything. Shopify currently runs roughly half a dozen of these distilled models in production, with more being added. Farhan also details River, Shopify's internal agentic substrate β€” a public-only Slack agent that queries their data warehouse, reads their PM system, and improves its own answers when engineers jump in to correct it. Plus: how Shopify governs AI-generated code at scale, why they killed their token leaderboard, how UCP is positioning Shopify's catalog for agentic commerce, and what a two-to-three year horizon looks like when agents start holding spending budgets and buying autonomously. πŸŽ™οΈ GUEST: Farhan Thawar | VP & Head of Engineering, Shopify πŸŽ™οΈ HOST: Sam Witteveen | VentureBeat __ If you enjoy these conversations, you ne… ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/beyond-the-pilot-enterprise-ai-in-action-7531113/episodes/small-models-massive-wins-the-new-shopify-ai-formula/transcription-requests` β€” Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/beyond-the-pilot-enterprise-ai-in-action-7531113/small-models-massive-wins-the-new-shopify-ai-formula.md` β€” Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.