Episode
The Hidden Cost of AI Model Inference at Scale
- Published
- Jul 2, 2026
- Duration seconds
- 528
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011/episodes/the-hidden-cost-of-ai-model-inference-at-scale/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011/the-hidden-cost-of-ai-model-inference-at-scale.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
Lucas and Luna unpack the often-overlooked expense of running AI models in production: inference costs. While training gets the headlines, inference is where the real spending happens for most companies. They explore how inference costs can dwarf training expenses, why hyperscalers are racing to build custom inference chips, and what the recent moves by companies like AMD and Nvidia reveal about the future of AI deployment. The hosts tie in the latest market data, including AMD's 1.6% weekly gain and the broader chip selloff, to show how investors are starting to price in the inference boom. If you're building an AI product or just trying to understand where the money goes in AI, this episode gives you a concrete framework for thinking about the cost of every query. #AIInference #InferenceCosts #AISpending #AMD #NVDA #ChipStocks #Hyperscalers #CustomSilicon #ModelDeployment #MachineLearning #AIChips #TechSpending #FexingoBusiness #BusinessPodcast #TechPodcast #AIEconomy #InferenceHardware #AICostModel Keep every episode free: buymeacoffee.com/fexingo