Episode

How SRE Teams Use Non-Abstract Large System Design to Prevent Outages

Podcast
The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering
Published
Jul 11, 2026
Duration seconds
514
Processing state
not_requested
Canonical source
https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0105.mp3
Audio
https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0105.mp3
JSON
/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-non-abstract-large-system-design-to-prevent-outages
Markdown
/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-non-abstract-large-system-design-to-prevent-outages.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-non-abstract-large-system-design-to-prevent-outages/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-non-abstract-large-system-design-to-prevent-outages.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

In Episode 105 of The Site Reliability Podcast with Fexingo, Lucas and Luna dive into Non-Abstract Large System Design (NALSD) — a method Google SRE teams use to evaluate distributed system designs before they hit production. They break down the six components of NALSD: system requirements, capacity estimation, failure modes, component interactions, data flow, and deployment architecture. Using a concrete example of a payment processing service with a five-nines uptime requirement, they show how NALSD helps catch design flaws early, reduce incident severity, and build shared mental models across teams. They also discuss how to integrate NALSD into sprint planning and post-incident reviews without adding overhead. Listeners will learn one specific technique to apply in their next design review. #SiteReliabilityEngineering #SRE #NonAbstractLargeSystemDesign #NALSD #ProductionEngineering #IncidentPrevention #SystemDesign #CapacityPlanning #FailureModeAnalysis #DistributedSystems #GoogleSRE #ReliabilityEngineering #Uptime #TechPodcast #FexingoBusiness #BusinessPodcast #Technology #Observability Keep every episode free: buymeacoffee.com/fexingo