# Training Machine Learning (ML) models on Kubernetes Page: https://stenobird.com/podcast/kubernetes-bytes-5263358/training-machine-learning-ml-models-on-kubernetes Text version: https://stenobird.com/podcast/kubernetes-bytes-5263358/training-machine-learning-ml-models-on-kubernetes.md Podcast: [Kubernetes Bytes](https://stenobird.com/podcast/kubernetes-bytes-5263358) Published: 2024-05-31T12:38:15+00:00 Episode link: https://zencastr.com/z/9OULupPH Audio file: https://redirect.zencastr.com/r/episode/6659c4b7cb4b86bda86ed868/size/79963037/audio-files/60f9ac743534330029a39c99/1cf2c760-d1b6-4f73-aaeb-895767010861.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/kubernetes-bytes-5263358/episodes/training-machine-learning-ml-models-on-kubernetes Duration seconds: 3329 ## Resource In this episode of the Kubernetes Bytes podcast, Bhavin sits down with Bernie Wu, VP Strategic Partnerships and AI/CXL/Kubernetes Initiatives at Memverge. They discuss about how Kubernetes is the most popular platform to run AI model training and model inferencing jobs. The discussion dives into model training, talking about different phases of a DAG, and then talk about how Memverge can help users with efficient and cost-effective model checkpoints. The discussion goes into topics like saving costs by using spot instances, hot restart of training jobs, reclaiming unused GPU resources, etc. Check out our website at https://kubernetesbytes.com/ [https://www.youtube.com/redirect?event=video_description&redir_token=QUFFLUhqbllTN0VpRHpuZWZiem9iWWc2bDF5NF9SMGhod3xBQ3Jtc0tuMnNGTmhwN01xdC01ZEFqRjFERnZBLUlUSTFCRWVta0tYQU4wWndITWhfV3RnbHh1cFVtbUc5NGZRandUcG1ocjQ5Q2R2cG1tQmpibkhJMVNRMHIycE1SYVJNdlVWMGFCa2xwTkhIOEFZUFRqVG5sdw&q=https%3A%2F%2Fkubernetesbytes.com%2F&v=Jf059PFn6l0] Episode Sponsor: Nethopper * Learn more about KAOPS: @nethopper.io * For a supported-demo: info@nethopper.io * Try the free version of KAOPS now! https://mynethopper.com/auth [https://www.youtube.com/redirect?event=video_description&redir_token=QUFFLUhqa2c2UWw4bFZISFN2X0hIWUtuWHlpWGk3WC1RZ3xBQ3Jtc0trQ3FpMlZDeXUwcVdZZGJ1d2R5RVREc1V0SzNPeXBzZGVCeWdzYjNsRndTVzRGcTYtUVFRUnI2bXEtWE96QVpNNThCaENiWXpqM3dzSmF6cml1c1BEYUk4bnlEQ3lqc1d0cXdTVWxuVE5YZWV3elNITQ&q=https%3A%2F%2Fmynethopper.com%2Fauth&v=Jf059PFn6l0] Cloud Native News: * https://www.aquasec.com/blog/linguistic-lumberjack-understanding-cve-2024-4323-in-fluent-bit/ * https://kubernetes.io/blog/2024/05/20/completing-cloud-provider-migration/ * https://thenewstack.io/introducing-aks-automatic-managed-kubernetes-for-developers/ * https://www.harness.io/blog/ha… ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/kubernetes-bytes-5263358/episodes/training-machine-learning-ml-models-on-kubernetes/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/kubernetes-bytes-5263358/training-machine-learning-ml-models-on-kubernetes.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.