Episode

Optimizing AI spend with Cast AI

Podcast
Kubernetes Bytes
Published
Aug 14, 2026
Duration seconds
3329
Processing state
not_requested
Canonical source
https://zencastr.com/z/npnYVeKd
Audio
https://redirect.zencastr.com/r/episode/6a7f178cb9a4cc9e347d8e16/size/79950540/audio-files/60f9ac743534330029a39c99/1e4a2cee-89a6-46d5-bce3-c9338828560e.mp3
JSON
/v1/public/podcasts/kubernetes-bytes-5263358/episodes/optimizing-ai-spend-with-cast-ai
Markdown
/podcast/kubernetes-bytes-5263358/optimizing-ai-spend-with-cast-ai.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/kubernetes-bytes-5263358/episodes/optimizing-ai-spend-with-cast-ai/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/kubernetes-bytes-5263358/optimizing-ai-spend-with-cast-ai.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

In this episode of the Kubernetes Bytes podcast, Bhavin talks to Phil Andrews, Global Field CTO at Cast AI about all things AI. The discussion starts by talking about why you don't need an Opus class or Fable/Mythos class model for each task, and how Kimchi from Cast AI can help manage the quality of your output with the cost associated with routing between different models. They also talk about Omni which allows customers to get GPUs across hyperscalers and neocloud providers. Listen to learn more! Check out our website at https://kubernetesbytes.com/ Show Notes: * https://www.linkedin.com/in/philip-andrews-iii/ * https://cast.ai/kimchi/ * https://docs.cast.ai/docs/omni-overview * https://cast.ai/blog/