{"podcast":{"title":"The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering","slug":"the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","podcast_index_feed_id":7871923,"rss_url":"https://feeds.fexingo.com/business/the-site-reliability-podcast.xml","website_url":"https://www.fexingo.com/","image_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/cover.png","author":"Fexingo","episode_count":119,"summary":"Lucas and Luna cut through the noise around site reliability engineering to examine how real-world SRE teams balance uptime, incident response, and production change. Each episode takes a single concept — error budgets, toil automation, postmortem culture, capacity planning — and grounds it in a specific case: how a major streaming service reduced paging noise, how a payments platform rebuilt its incident command structure, or how a cloud provider manages multi-region failover. Lucas brings the numbers — latency percentiles, MTTR trends, SLO burn rates — while Luna pushes on the human and organizational trade-offs: What does a junior SRE need to know about on-call? How do you measure reliability without crushing innovation? Why do some blameless postmortems actually work? Together they treat SRE not as a certification topic but as a living practice, citing real outages, open-source tools, and engineering blogs. This show is for engineers, ops leads, and platform teams who already know the basics and want to debate the hard edges: Is 99.999% uptime always worth the cost? When should you deliberately degrade service to improve reliability? How do you design for resilience when your…","last_synced_at":"2026-07-18T22:18:41.635042+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923"},"episode":{"title":"How SRE Teams Use Observability Pipelines to Reduce Data Costs","slug":"how-sre-teams-use-observability-pipelines-to-reduce-data-costs","published_at":"2026-07-08T12:09:31+00:00","page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-observability-pipelines-to-reduce-data-costs","show_page_url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923","url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0098.mp3","audio_url":"https://audio.fexingo.com/business/the-site-reliability-podcast/episode-0098.mp3","summary":"Episode 98 of The Site Reliability Podcast with Fexingo dives into observability pipelines — the middleware that sits between your systems and your monitoring tools. Lucas and Luna explore how teams at companies like DoorDash and Slack have used pipelines to filter, sample, and shape telemetry data before it reaches their observability backends, slashing costs by up to 70 percent without losing signal. They break down the concept of 'intelligent sampling' versus naive tail-based sampling, discuss the trade-offs of using open-source tools like OpenTelemetry Collector versus managed services like Cribl or Vector, and walk through a concrete example: how a mid-stage SaaS company reduced its Datadog bill from $240,000 to $72,000 per year. The episode also covers the pitfalls — like accidentally dropping critical traces during partial outages — and how to design pipeline rules that preserve high-cardinality dimensions for debugging. If you're an SRE or platform engineer feeling the sting of observability vendor pricing, this one's for you. #ObservabilityPipelines #SRE #SiteReliabilityEngineering #DataCostOptimization #OpenTelemetry #Cribl #Vector #Datadog #DoorDash #Slack #TelemetrySampling #IntelligentSampling #ObservabilityCosts #Monitoring #Technology #FexingoBusiness #BusinessPodcast #ProdEng Keep every episode free: buymeacoffee.com/fexingo","meta_description":"Episode 98 of The Site Reliability Podcast with Fexingo dives into observability pipelines — the middleware that sits between your systems and your monito…","key_points":[],"chapters":[],"topics":[],"duration_seconds":668,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/episodes/how-sre-teams-use-observability-pipelines-to-reduce-data-costs/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/the-site-reliability-podcast-with-fexingo-sre-uptime-and-production-engineering-7871923/how-sre-teams-use-observability-pipelines-to-reduce-data-costs.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}