{"podcast":{"title":"Kubernetes Podcast from Google","slug":"kubernetes-podcast-from-google","podcast_index_feed_id":860803,"rss_url":"https://rss.libsyn.com/shows/419861/destinations/3486674.xml","website_url":"https://kubernetespodcast.com","image_url":"https://static.libsyn.com/p/assets/a/3/a/a/a3aa4f08236059f5e55e3c100dce7605/NewKPodRoboWhite-20240924-1zwoxpn1sl.png","author":"Kubernetes Podcast from Google","episode_count":265,"summary":"A biweekly podcast focused on what's happening in the Kubernetes community hosted by Abdel Sghiouar and Kaslin Fields. We cover Kubernetes, cloud-native applications, and other developments in the ecosystem. Abdel and Kaslin on Twitter at @KubernetesPod or by email at kubernetespodcast@google.com.","last_synced_at":null,"page_url":"https://stenobird.com/podcast/kubernetes-podcast-from-google"},"episode":{"title":"HPC Workload Scheduling, with Ricardo Rocha","slug":"hpc-workload-scheduling-with-ricardo-rocha","published_at":"2025-07-09T00:46:00+00:00","page_url":"https://stenobird.com/podcast/kubernetes-podcast-from-google/hpc-workload-scheduling-with-ricardo-rocha","show_page_url":"https://stenobird.com/podcast/kubernetes-podcast-from-google","url":"https://e780d51f-f115-44a6-8252-aed9216bb521.libsyn.com/hpc-workload-scheduling-with-ricardo-rocha","audio_url":"https://traffic.libsyn.com/secure/e780d51f-f115-44a6-8252-aed9216bb521/KPOD255.mp3?dest-id=3486674","summary":"Ricardo Rocha from CERN explains how to bridge the gap between traditional High-Performance Computing (HPC) and cloud-native Kubernetes environments. The discussion focuses on using specialized schedulers like Kueue to manage expensive, high-demand resources like GPUs without reinventing the Kubernetes core.","meta_description":"Learn how CERN uses Kubernetes for HPC workloads, managing expensive GPU resources, and the importance of extending rather than replacing Kubernetes core.","key_points":["Main idea: Kubernetes is evolving from a stateless web-app orchestrator to a platform capable of handling complex, resource-intensive scientific workloads","Practical takeaway: Use 'out-of-tree' controllers like Kueue to implement specialized logic (like pod suspension) while leveraging native Kubernetes primitives to avoid maintenance debt","Failure mode: Building parallel execution systems that bypass Kubernetes semantics leads to massive technical debt as you must manually reimplement every core upstream improvement","Main idea: The rise of expensive, pre-committed GPU instances is shifting cloud computing from an 'on-demand' model back toward a 'reservation' model similar to on-premises hardware","Practical takeaway: Focus on extending Kubernetes via the Gateway API and specialized controllers rather than fighting the existing system's architecture"],"chapters":[{"start_ms":60000,"title":"Node Feature Discovery and Gemini CLI","summary":"An introduction to NFD for hardware-aware scheduling and the new Gemini CLI for interacting with AI from the terminal."},{"start_ms":245000,"title":"CERN's Resource Challenges","summary":"Discussing the necessity of optimizing fixed budgets and managing massive datasets in scientific computing."},{"start_ms":440000,"title":"Low-Level Optimization in HPC","summary":"The shift toward needing CPU pinning, NUMA awareness, and high-efficiency node usage for research workloads."},{"start_ms":625000,"title":"The Evolution of Kubernetes for Research","summary":"Why it took time for scientific and research-oriented voices to integrate into the Kubernetes ecosystem."},{"start_ms":830000,"title":"Managing Expensive GPU Resources","summary":"The difficulty of managing heterogeneous hardware and preventing over-provisioning of high-cost accelerators."},{"start_ms":1020000,"title":"The Role of Specialized Schedulers","summary":"A look at the history of grid computing and how projects like Volcano and UniKorn provide necessary extensions."},{"start_ms":1205000,"title":"The CNCF Ecosystem and Maturity","summary":"How the CNCF sandbox and incubation process supports the growth of specialized cloud-native projects."},{"start_ms":1400000,"title":"Future of Batch Workloads","summary":"Speculating on multi-cluster, multi-region management and the move toward resource reservation models."}],"topics":["Kubernetes","HPC","CERN","GPU Scheduling","Kueue","Cloud Native","CNCF","Batch Workloads","Infrastructure Engineering"],"duration_seconds":2578,"processing_state":"processed","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/kubernetes-podcast-from-google/episodes/hpc-workload-scheduling-with-ricardo-rocha/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/kubernetes-podcast-from-google/hpc-workload-scheduling-with-ricardo-rocha.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}