# An Agentic Mixture of Experts for DevOps with Sunil Mallya - #708

Page: https://stenobird.com/podcast/twiml-ai-podcast/an-agentic-mixture-of-experts-for-devops-with-sunil-mallya-708
Text version: https://stenobird.com/podcast/twiml-ai-podcast/an-agentic-mixture-of-experts-for-devops-with-sunil-mallya-708.md
Podcast: [The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)](https://stenobird.com/podcast/twiml-ai-podcast)
Published: 2024-11-04T13:53:00+00:00
Episode link: https://twimlai.com/podcast/twimlai/an-agentic-mixture-of-experts-for-devops/
Audio file: https://pscrb.fm/rss/p/traffic.megaphone.fm/MLN8491913296.mp3?updated=1730753189
Processing state: failed
JSON: https://stenobird.com/v1/public/podcasts/twiml-ai-podcast/episodes/an-agentic-mixture-of-experts-for-devops-with-sunil-mallya-708
Duration seconds: 4509

## Resource

Today we're joined by Sunil Mallya, CTO and co-founder of Flip AI. We discuss Flip’s incident debugging system for DevOps, which was built using a custom mixture of experts (MoE) large language model (LLM) trained on a novel "CoMELT" observability dataset which combines traditional MELT data—metrics, events, logs, and traces—with code to efficiently identify root failure causes in complex software systems. We discuss the challenges of integrating time-series data with LLMs and their multi-decoder architecture designed for this purpose. Sunil describes their system's agent-based design, focusing on clear roles and boundaries to ensure reliability. We examine their "chaos gym," a reinforcement learning environment used for testing and improving the system's robustness. Finally, we discuss the practical considerations of deploying such a system at scale in diverse environments and much more. The complete show notes for this episode can be found at https://twimlai.com/go/708.

## Actions

- request_transcript: `POST https://stenobird.com/v1/public/podcasts/twiml-ai-podcast/episodes/an-agentic-mixture-of-experts-for-devops-with-sunil-mallya-708/transcription-requests` — Idempotently request low-priority transcript generation for this episode.
- read_markdown: `GET https://stenobird.com/podcast/twiml-ai-podcast/an-agentic-mixture-of-experts-for-devops-with-sunil-mallya-708.md` — Read the agent-friendly Markdown representation of this episode resource.

A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed.

## Transcript

Full transcripts are not published on public pages unless there is a clear rights basis.