Episode
How Data Scientists Use Distributed Computing for Massive Datasets
- Podcast
- The Data Science Podcast with Fexingo: Analytics, Machine Learning, and Data-Driven Conversations
- Published
- Jun 13, 2026
- Duration seconds
- 481
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-distributed-computing-for-massive-datasets/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-distributed-computing-for-massive-datasets.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
When your dataset outgrows a single machine, what do you do? In this episode, Lucas and Luna explore how data scientists use distributed computing frameworks like Apache Spark and Dask to process terabytes of data without crashing their laptops. They break down the key concept of data partitioning, explain why MapReduce is still relevant, and walk through a real example of how a mid-sized e-commerce company reorganized its log-processing pipeline to cut runtime from 14 hours to 47 minutes. Lucas shares a cautionary tale about shuffling bottlenecks that can ruin a cluster's performance, and Luna asks the practical question every team faces: when does it make sense to move from a single-node pandas workflow to a distributed system? They also discuss managed services like Databricks and AWS EMR versus rolling your own cluster. No prior distributed systems experience required — just a curiosity about what happens when data gets too big for a spreadsheet. #DataScience #DistributedComputing #ApacheSpark #Dask #MapReduce #BigData #DataEngineering #DataPartitioning #Shuffling #Databricks #AWSEmr #Pandas #Tech #Technology #FexingoBusiness #BusinessPodcast #Podcast #DataPodcast Keep every episode free: buymeacoffee.com/fexingo