# The Hidden Cost of AI Model Inference at Scale Page: https://stenobird.com/podcast/the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011/the-hidden-cost-of-ai-model-inference-at-scale Text version: https://stenobird.com/podcast/the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011/the-hidden-cost-of-ai-model-inference-at-scale.md Podcast: [The AI Podcast with Fexingo: Artificial Intelligence, Machine Learning, and Modern AI Models](https://stenobird.com/podcast/the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011) Published: 2026-07-02T08:11:05+00:00 Episode link: https://audio.fexingo.com/business/the-ai-podcast/episode-0086.mp3 Audio file: https://audio.fexingo.com/business/the-ai-podcast/episode-0086.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011/episodes/the-hidden-cost-of-ai-model-inference-at-scale Duration seconds: 528 ## Resource Lucas and Luna unpack the often-overlooked expense of running AI models in production: inference costs. While training gets the headlines, inference is where the real spending happens for most companies. They explore how inference costs can dwarf training expenses, why hyperscalers are racing to build custom inference chips, and what the recent moves by companies like AMD and Nvidia reveal about the future of AI deployment. The hosts tie in the latest market data, including AMD's 1.6% weekly gain and the broader chip selloff, to show how investors are starting to price in the inference boom. If you're building an AI product or just trying to understand where the money goes in AI, this episode gives you a concrete framework for thinking about the cost of every query. #AIInference #InferenceCosts #AISpending #AMD #NVDA #ChipStocks #Hyperscalers #CustomSilicon #ModelDeployment #MachineLearning #AIChips #TechSpending #FexingoBusiness #BusinessPodcast #TechPodcast #AIEconomy #InferenceHardware #AICostModel Keep every episode free: buymeacoffee.com/fexingo ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011/episodes/the-hidden-cost-of-ai-model-inference-at-scale/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011/the-hidden-cost-of-ai-model-inference-at-scale.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.