# How Cloud Bills Are Adding AI Inference Surcharges Page: https://stenobird.com/podcast/cloud-computing-with-fexingo-aws-azure-gcp-and-modern-infrastructure-conversations-7871918/how-cloud-bills-are-adding-ai-inference-surcharges Text version: https://stenobird.com/podcast/cloud-computing-with-fexingo-aws-azure-gcp-and-modern-infrastructure-conversations-7871918/how-cloud-bills-are-adding-ai-inference-surcharges.md Podcast: [Cloud Computing with Fexingo: AWS, Azure, GCP, and Modern Infrastructure Conversations](https://stenobird.com/podcast/cloud-computing-with-fexingo-aws-azure-gcp-and-modern-infrastructure-conversations-7871918) Published: 2026-06-26T08:20:13+00:00 Episode link: https://audio.fexingo.com/business/cloud-computing/episode-0074.mp3 Audio file: https://audio.fexingo.com/business/cloud-computing/episode-0074.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/cloud-computing-with-fexingo-aws-azure-gcp-and-modern-infrastructure-conversations-7871918/episodes/how-cloud-bills-are-adding-ai-inference-surcharges Duration seconds: 412 ## Resource In this episode, Lucas and Luna dig into a new line item showing up on enterprise cloud invoices: the AI inference surcharge. Amazon, Microsoft, and Google are now charging a premium per million tokens when customers use their managed inference APIs on AWS Bedrock, Azure OpenAI Service, and Vertex AI. The hosts break down why the premium ranges from 30% to 120% above base compute costs, how it's tied to NVIDIA's H100 GPU scarcity and the cost of high-bandwidth memory, and what it means for startups building AI features. They also discuss the fine print: some providers waive the surcharge if you commit to reserved GPU instances, while others levy it even on spot usage. Real numbers from a real bill show a 47% increase in monthly spend for a mid-stage startup. This episode is a practical guide to understanding and negotiating the newest cloud cost. #AIInference #CloudCosts #AWSBedrock #AzureOpenAI #VertexAI #NVIDIAH100 #GPUScarcity #EnterpriseBilling #CloudPricing #Technology #TechPodcast #FexingoBusiness #BusinessPodcast #LucasAndLuna #CloudComputing #InferenceSurcharge #TokenPricing #StartupCosts Keep every episode free: buymeacoffee.com/fexingo ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/cloud-computing-with-fexingo-aws-azure-gcp-and-modern-infrastructure-conversations-7871918/episodes/how-cloud-bills-are-adding-ai-inference-surcharges/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/cloud-computing-with-fexingo-aws-azure-gcp-and-modern-infrastructure-conversations-7871918/how-cloud-bills-are-adding-ai-inference-surcharges.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.