Episode

How Intercom Cut $250K/Month by Ditching GPT for Qwen

Podcast
Chain of Thought | AI Agents, Infrastructure & Engineering
Published
Feb 26, 2026
Duration seconds
3211
Processing state
not_requested
Canonical source
https://share.transistor.fm/s/0fd18337
Audio
https://media.transistor.fm/0fd18337/b333c541.mp3
JSON
/v1/public/podcasts/chain-of-thought-ai-agents/episodes/how-intercom-cut-250k-month-by-ditching-gpt-for-qwen
Markdown
/podcast/chain-of-thought-ai-agents/how-intercom-cut-250k-month-by-ditching-gpt-for-qwen.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/chain-of-thought-ai-agents/episodes/how-intercom-cut-250k-month-by-ditching-gpt-for-qwen/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/chain-of-thought-ai-agents/how-intercom-cut-250k-month-by-ditching-gpt-for-qwen.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

Intercom was spending $250K/month on a single summarization task using GPT. Then they replaced it with a fine-tuned 14B parameter Qwen model and saved almost all of it. In this episode, Intercom's Chief AI Officer, Fergal Reid, walks through exactly how they made that call, where their approach has changed over time, and how all of their efforts built their Fin customer service agent. Chain of Thought is hosted by Conor Bronsdon.