# Architecting Kubernetes for GPU-Accelerated AI Applications Page: https://stenobird.com/podcast/the-business-compass-llc-podcasts-7078188/architecting-kubernetes-for-gpu-accelerated-ai-applications Text version: https://stenobird.com/podcast/the-business-compass-llc-podcasts-7078188/architecting-kubernetes-for-gpu-accelerated-ai-applications.md Podcast: [The Business Compass LLC Podcasts](https://stenobird.com/podcast/the-business-compass-llc-podcasts-7078188) Published: 2026-06-17T05:02:06+00:00 Episode link: https://podcast.businesscompassllc.com/e/architecting-kubernetes-for-gpu-accelerated-ai-applications/ Audio file: https://mcdn.podbean.com/mf/web/2m76szsrz4rn4pxb/f6a18d70-3444-42a6-b4b3-34a04cfedb35.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/the-business-compass-llc-podcasts-7078188/episodes/architecting-kubernetes-for-gpu-accelerated-ai-applications Duration seconds: 1787 ## Resource Running AI workloads at scale is hard. Running them efficiently on Kubernetes without wasting expensive GPU resources is even harder. If you’re a platform engineer, ML engineer, or DevOps architect trying to get serious about GPU cluster management for AI, this guide is built for you. ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/the-business-compass-llc-podcasts-7078188/episodes/architecting-kubernetes-for-gpu-accelerated-ai-applications/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/the-business-compass-llc-podcasts-7078188/architecting-kubernetes-for-gpu-accelerated-ai-applications.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.