# How Data Scientists Use Distributed Computing for Massive Datasets Page: https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-distributed-computing-for-massive-datasets Text version: https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-distributed-computing-for-massive-datasets.md Podcast: [The Data Science Podcast with Fexingo: Analytics, Machine Learning, and Data-Driven Conversations](https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831) Published: 2026-06-13T21:07:20+00:00 Episode link: https://audio.fexingo.com/business/the-data-science-podcast/episode-0049.mp3 Audio file: https://audio.fexingo.com/business/the-data-science-podcast/episode-0049.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-distributed-computing-for-massive-datasets Duration seconds: 481 ## Resource When your dataset outgrows a single machine, what do you do? In this episode, Lucas and Luna explore how data scientists use distributed computing frameworks like Apache Spark and Dask to process terabytes of data without crashing their laptops. They break down the key concept of data partitioning, explain why MapReduce is still relevant, and walk through a real example of how a mid-sized e-commerce company reorganized its log-processing pipeline to cut runtime from 14 hours to 47 minutes. Lucas shares a cautionary tale about shuffling bottlenecks that can ruin a cluster's performance, and Luna asks the practical question every team faces: when does it make sense to move from a single-node pandas workflow to a distributed system? They also discuss managed services like Databricks and AWS EMR versus rolling your own cluster. No prior distributed systems experience required — just a curiosity about what happens when data gets too big for a spreadsheet. #DataScience #DistributedComputing #ApacheSpark #Dask #MapReduce #BigData #DataEngineering #DataPartitioning #Shuffling #Databricks #AWSEmr #Pandas #Tech #Technology #FexingoBusiness #BusinessPodcast #Podcast #DataPodcast Keep every episode free: buymeacoffee.com/fexingo ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-distributed-computing-for-massive-datasets/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-distributed-computing-for-massive-datasets.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.