# Why Cloud Providers Now Bill by the Millisecond for GPUs Page: https://stenobird.com/podcast/cloud-computing-with-fexingo-aws-azure-gcp-and-modern-infrastructure-conversations-7871918/why-cloud-providers-now-bill-by-the-millisecond-for-gpus Text version: https://stenobird.com/podcast/cloud-computing-with-fexingo-aws-azure-gcp-and-modern-infrastructure-conversations-7871918/why-cloud-providers-now-bill-by-the-millisecond-for-gpus.md Podcast: [Cloud Computing with Fexingo: AWS, Azure, GCP, and Modern Infrastructure Conversations](https://stenobird.com/podcast/cloud-computing-with-fexingo-aws-azure-gcp-and-modern-infrastructure-conversations-7871918) Published: 2026-06-23T08:18:25+00:00 Episode link: https://audio.fexingo.com/business/cloud-computing/episode-0068.mp3 Audio file: https://audio.fexingo.com/business/cloud-computing/episode-0068.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/cloud-computing-with-fexingo-aws-azure-gcp-and-modern-infrastructure-conversations-7871918/episodes/why-cloud-providers-now-bill-by-the-millisecond-for-gpus Duration seconds: 560 ## Resource In episode 68 of Cloud Computing with Fexingo, Lucas and Luna dive into a quiet-but-radical shift in cloud pricing: GPU instances billed by the millisecond instead of the second or hour. They trace the change from NVIDIA's H100 to the upcoming Rubin architecture, explain why providers like AWS, Azure, and Google Cloud are moving to sub-second billing for AI workloads, and unpack what this means for inference costs, bursty training jobs, and enterprise budgeting. Along the way, they cite specific examples—like a 47% cost reduction for a 500-millisecond inference call on an H100 cluster—and discuss how the shift pressures traditional cloud financial management tools. If you're running AI workloads or managing cloud spend, this episode gives you one concrete number and one negotiation tactic to use in your next contract renewal. #CloudComputing #GPU #AIBilling #MillisecondPricing #AWS #Azure #GoogleCloud #H100 #Rubin #NVIDIA #InferenceCosts #CloudEconomics #FinOps #Technology #BusinessPodcast #FexingoBusiness #CloudInfrastructure #AITraining Keep every episode free: buymeacoffee.com/fexingo ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/cloud-computing-with-fexingo-aws-azure-gcp-and-modern-infrastructure-conversations-7871918/episodes/why-cloud-providers-now-bill-by-the-millisecond-for-gpus/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/cloud-computing-with-fexingo-aws-azure-gcp-and-modern-infrastructure-conversations-7871918/why-cloud-providers-now-bill-by-the-millisecond-for-gpus.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.