# How Pinterest Rebuilt Its Recommendation Engine for 500 Million Users Page: https://stenobird.com/podcast/the-cto-podcast-with-fexingo-technical-leadership-architecture-and-engineering-org-7871807/how-pinterest-rebuilt-its-recommendation-engine-for-500-million-users Text version: https://stenobird.com/podcast/the-cto-podcast-with-fexingo-technical-leadership-architecture-and-engineering-org-7871807/how-pinterest-rebuilt-its-recommendation-engine-for-500-million-users.md Podcast: [The CTO Podcast with Fexingo: Technical Leadership, Architecture, and Engineering Org](https://stenobird.com/podcast/the-cto-podcast-with-fexingo-technical-leadership-architecture-and-engineering-org-7871807) Published: 2026-06-20T19:40:54+00:00 Episode link: https://audio.fexingo.com/business/the-cto-podcast/episode-0063.mp3 Audio file: https://audio.fexingo.com/business/the-cto-podcast/episode-0063.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/the-cto-podcast-with-fexingo-technical-leadership-architecture-and-engineering-org-7871807/episodes/how-pinterest-rebuilt-its-recommendation-engine-for-500-million-users Duration seconds: 824 ## Resource In this episode of The CTO Podcast, Lucas and Luna dive into how Pinterest's engineering team rebuilt its core recommendation engine from a batch-processing pipeline to a real-time, graph-based system serving over 500 million monthly active users. They explore the specific architectural decisions Pinterest made: moving from collaborative filtering to a heterogeneous graph neural network called PinSage, deploying it on TensorFlow Serving with Kubernetes for low-latency inference, and handling the cold-start problem for new pins and users. The discussion covers the trade-offs between offline batch precomputation and online inference, how Pinterest reduced recommendation latency from hours to milliseconds, and the infrastructure costs involved. Lucas explains the graph-based approach that captures user intent through 'pins' and 'boards,' while Luna questions how this impacts content discovery and engagement. Tune in for a detailed look at a real-world scale challenge in modern machine learning systems. #Pinterest #RecommendationEngine #GraphNeuralNetworks #PinSage #MachineLearning #TensorFlow #Kubernetes #RealTimeInference #ColdStart #ContentDiscovery #EngineeringArchitecture #Scale #BusinessAndTechnology #FexingoBusiness #BusinessPodcast #CTO #TechnicalLeadership #MLInfrastructure Keep every episode free: buymeacoffee.com/fexingo ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/the-cto-podcast-with-fexingo-technical-leadership-architecture-and-engineering-org-7871807/episodes/how-pinterest-rebuilt-its-recommendation-engine-for-500-million-users/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/the-cto-podcast-with-fexingo-technical-leadership-architecture-and-engineering-org-7871807/how-pinterest-rebuilt-its-recommendation-engine-for-500-million-users.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.