Episode
How Datadog Monitors Its Own Infrastructure
- Published
- Jun 18, 2026
- Duration seconds
- 493
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/the-cto-podcast-with-fexingo-technical-leadership-architecture-and-engineering-org-7871807/episodes/how-datadog-monitors-its-own-infrastructure/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/the-cto-podcast-with-fexingo-technical-leadership-architecture-and-engineering-org-7871807/how-datadog-monitors-its-own-infrastructure.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
Episode 58 of The CTO Podcast goes inside Datadog's engineering org to explore how the company monitors its own 100-terabyte infrastructure. Lucas and Luna walk through Datadog's dogfooding culture, the architectural challenges of running a monitoring platform for itself, and how the team handles alert fatigue, distributed tracing, and log ingestion at massive scale. They discuss specific tools like the Datadog Agent, the trace-agent, and the custom time-series database built in-house. The episode includes concrete numbers: 30 trillion time-series points ingested daily, 99.99 percent uptime target, and how the SRE team manages 8,000 hosts across multiple cloud providers. Tune in for a rare look at how the watcher watches itself. #Datadog #InfrastructureMonitoring #Dogfooding #SRE #Observability #TimeSeriesDatabase #DistributedTracing #AlertFatigue #CloudInfrastructure #EngineeringCulture #SiteReliabilityEngineering #DevOps #BusinessAndTechnology #FexingoBusiness #BusinessPodcast #CTO #TechnicalLeadership #Architecture Keep every episode free: buymeacoffee.com/fexingo