Episode
How Datadog Monitors Its Own 100-Terabyte Infrastructure
- Published
- Jun 16, 2026
- Duration seconds
- 595
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/the-cto-podcast-with-fexingo-technical-leadership-architecture-and-engineering-org-7871807/episodes/how-datadog-monitors-its-own-100-terabyte-infrastructure/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/the-cto-podcast-with-fexingo-technical-leadership-architecture-and-engineering-org-7871807/how-datadog-monitors-its-own-100-terabyte-infrastructure.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
Episode 54 of The CTO Podcast: Lucas and Luna explore how Datadog, the monitoring giant, uses its own tools to manage a sprawling infrastructure that ingests over 100 terabytes of data daily. They dive into the dogfooding strategy, the architectural choices that keep observability scalable, and the surprising insight that Datadog runs its entire backend on a single PostgreSQL fork — with custom sharding. Lucas explains the engineering org structure behind the monitoring team, and Luna questions whether dogfooding can blind teams to customer pain. Specific examples include how Datadog handles metric cardinality explosion and why they built a separate time-series database internally before launching it as a product. #Datadog #Observability #Dogfooding #TechLeadership #Infrastructure #PostgreSQL #Scalability #TimeSeriesDatabase #EngineeringCulture #Monitoring #CTOPodcast #FexingoBusiness #BusinessPodcast #Architecture #Sharding #MetricCardinality #SRE #CloudNative Keep every episode free: buymeacoffee.com/fexingo