{"podcast":{"title":"Kubernetes Podcast from Google","slug":"kubernetes-podcast-from-google","podcast_index_feed_id":860803,"rss_url":"https://rss.libsyn.com/shows/419861/destinations/3486674.xml","website_url":"https://kubernetespodcast.com","image_url":"https://static.libsyn.com/p/assets/a/3/a/a/a3aa4f08236059f5e55e3c100dce7605/NewKPodRoboWhite-20240924-1zwoxpn1sl.png","author":"Kubernetes Podcast from Google","episode_count":265,"summary":"A biweekly podcast focused on what's happening in the Kubernetes community hosted by Abdel Sghiouar and Kaslin Fields. We cover Kubernetes, cloud-native applications, and other developments in the ecosystem. Abdel and Kaslin on Twitter at @KubernetesPod or by email at kubernetespodcast@google.com.","last_synced_at":null,"page_url":"https://stenobird.com/podcast/kubernetes-podcast-from-google"},"episode":{"title":"Multi-Cluster Orchestrator, with Nick Eberts and Jon Li","slug":"multi-cluster-orchestrator-with-nick-eberts-and-jon-li","published_at":"2025-05-28T15:21:00+00:00","page_url":"https://stenobird.com/podcast/kubernetes-podcast-from-google/multi-cluster-orchestrator-with-nick-eberts-and-jon-li","show_page_url":"https://stenobird.com/podcast/kubernetes-podcast-from-google","url":"https://e780d51f-f115-44a6-8252-aed9216bb521.libsyn.com/multi-cluster-orchestrator-with-nick-eberts-and-jon-li","audio_url":"https://traffic.libsyn.com/secure/e780d51f-f115-44a6-8252-aed9216bb521/KPOD253.mp3?dest-id=3486674","summary":"The Multi-Cluster Orchestrator (MCO) solves the challenge of managing workloads across fragmented, non-uniform cloud resources. It provides a standardized way to scale AI inference engines from zero to one across multiple clusters based on real-time capacity and hardware availability.","meta_description":"Learn how the new Multi-Cluster Orchestrator (MCO) manages AI inference workloads and handles hardware scarcity across distributed Kubernetes clusters.","key_points":["Main idea: MCO addresses the shift from 'infinite' cloud capacity to a reality of hardware stockouts and non-uniform GPU availability","Practical takeaway: Use MCO to automate workload placement recommendations, allowing tools like Argo CD to scale workloads from zero to one in available regions","Technical detail: The system utilizes a normalized 'Cluster Profile' spec to unify cluster inventories across different providers and tools","Failure mode: Outdated cluster profiles can lead to MCO making recommendations based on stale capacity data","Practical takeaway: For AI inference, leverage custom metrics like KV cache utilization or queue depth to drive precise scaling decisions"],"chapters":[{"start_ms":60000,"title":"Kubernetes & GKE Updates","summary":"Brief overview of Kubernetes 1.33 availability in GKE Rapid channel and Kyverno 1.14.0 release."},{"start_ms":155000,"title":"Introduction to MCO","summary":"An introduction to the Multi-Cluster Orchestrator and the shift in cloud computing assumptions."},{"start_ms":245000,"title":"The Problem: Hardware Scarcity","summary":"Discussing why the assumption of infinite, uniform cloud capacity no longer holds true for modern workloads."},{"start_ms":340000,"title":"Defining Multi-Cluster Orchestrator","summary":"Explaining how MCO manages workloads across multiple clusters to optimize for proximity and availability."},{"start_ms":435000,"title":"Scaling AI Inference","summary":"How MCO handles the 'zero to one' scaling problem for compute-intensive inference engines."},{"start_ms":525000,"title":"Capacity and Metrics","summary":"Using open, accessible metrics to determine regional capacity and drive workload placement."},{"start_ms":615000,"title":"Latency in LLM Inference","summary":"The challenges of managing high-latency, multi-modal LLM requests compared to traditional microservices."},{"start_ms":715000,"title":"Standardizing Cluster Inventories","summary":"Using Cluster Profiles to create a normalized, open-source specification for multi-cluster management."}],"topics":["Kubernetes","Multi-Cluster Orchestrator","AI Inference","GKE","Cloud Infrastructure","GPU Availability","Cluster Management","LLM Deployment"],"duration_seconds":1291,"processing_state":"processed","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/kubernetes-podcast-from-google/episodes/multi-cluster-orchestrator-with-nick-eberts-and-jon-li/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/kubernetes-podcast-from-google/multi-cluster-orchestrator-with-nick-eberts-and-jon-li.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}