{"podcast":{"title":"Data Science Tech Brief By HackerNoon","slug":"data-science-tech-brief-by-hackernoon-6367564","podcast_index_feed_id":6367564,"rss_url":"https://feeds.transistor.fm/data-science-tech-brief-by-hackernoon","website_url":null,"image_url":"https://img.transistorcdn.com/PRg81mb1bHdu71bs3zSzRC6oEjt9WcIHjS2ba3uMWCY/rs:fill:0:0:1/w:1400/h:1400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS9zaG93/LzQxMjY4LzE2ODM1/ODI1ODUtYXJ0d29y/ay5qcGc.jpg","author":"HackerNoon","episode_count":100,"summary":"Learn the latest data science updates in the tech world.","last_synced_at":"2026-06-18T06:17:56.540839+00:00","page_url":"https://stenobird.com/podcast/data-science-tech-brief-by-hackernoon-6367564"},"episode":{"title":"Optimizing Distributed Data Processing for ML at Scale","slug":"optimizing-distributed-data-processing-for-ml-at-scale","published_at":"2026-05-21T16:00:43+00:00","page_url":"https://stenobird.com/podcast/data-science-tech-brief-by-hackernoon-6367564/optimizing-distributed-data-processing-for-ml-at-scale","show_page_url":"https://stenobird.com/podcast/data-science-tech-brief-by-hackernoon-6367564","url":"https://share.transistor.fm/s/993507ed","audio_url":"https://media.transistor.fm/993507ed/2a432a6d.mp3","summary":"This story was originally published on HackerNoon at: https://hackernoon.com/optimizing-distributed-data-processing-for-ml-at-scale . A practitioner's guide to ML data pipeline performance: read the query plan first, eliminate shuffle, fix file layout, handle skew, prune columns Check more stories related to data-science at: https://hackernoon.com/c/data-science . You can also check exclusive content about #spark , #pyspark , #machine-learning , #data-engineering , #performance-optimization , #distributed-systems , #distributed-data-processing , #optimizing-distributed-data , and more. This story was written by: @seshendranath . Learn more about this writer by checking @seshendranath's about page, and for more stories, please visit hackernoon.com . Stop tuning knobs on a broken foundation shuffle, file layout, skew, and column pruning do more for ML pipeline performance than any clever algorithm.","meta_description":"This story was originally published on HackerNoon at: https://hackernoon.com/optimizing-distributed-data-processing-for-ml-at-scale . A practitioner's gui…","key_points":[],"chapters":[],"topics":[],"duration_seconds":423,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/data-science-tech-brief-by-hackernoon-6367564/episodes/optimizing-distributed-data-processing-for-ml-at-scale/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/data-science-tech-brief-by-hackernoon-6367564/optimizing-distributed-data-processing-for-ml-at-scale.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}