Episode

How Data Scientists Use Distributed Computing for Massive Datasets

Podcast
The Data Science Podcast with Fexingo: Analytics, Machine Learning, and Data-Driven Conversations
Published
Jun 13, 2026
Duration seconds
481
Processing state
not_requested
Canonical source
https://audio.fexingo.com/business/the-data-science-podcast/episode-0049.mp3
Audio
https://audio.fexingo.com/business/the-data-science-podcast/episode-0049.mp3
JSON
/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-distributed-computing-for-massive-datasets
Markdown
/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-distributed-computing-for-massive-datasets.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-distributed-computing-for-massive-datasets/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-distributed-computing-for-massive-datasets.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

When your dataset outgrows a single machine, what do you do? In this episode, Lucas and Luna explore how data scientists use distributed computing frameworks like Apache Spark and Dask to process terabytes of data without crashing their laptops. They break down the key concept of data partitioning, explain why MapReduce is still relevant, and walk through a real example of how a mid-sized e-commerce company reorganized its log-processing pipeline to cut runtime from 14 hours to 47 minutes. Lucas shares a cautionary tale about shuffling bottlenecks that can ruin a cluster's performance, and Luna asks the practical question every team faces: when does it make sense to move from a single-node pandas workflow to a distributed system? They also discuss managed services like Databricks and AWS EMR versus rolling your own cluster. No prior distributed systems experience required — just a curiosity about what happens when data gets too big for a spreadsheet. #DataScience #DistributedComputing #ApacheSpark #Dask #MapReduce #BigData #DataEngineering #DataPartitioning #Shuffling #Databricks #AWSEmr #Pandas #Tech #Technology #FexingoBusiness #BusinessPodcast #Podcast #DataPodcast Keep every episode free: buymeacoffee.com/fexingo