Episode

How Data Scientists Use Data Version Control for Reproducibility

Podcast
The Data Science Podcast with Fexingo: Analytics, Machine Learning, and Data-Driven Conversations
Published
Jul 13, 2026
Duration seconds
756
Processing state
not_requested
Canonical source
https://audio.fexingo.com/business/the-data-science-podcast/episode-0107.mp3
Audio
https://audio.fexingo.com/business/the-data-science-podcast/episode-0107.mp3
JSON
/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-data-version-control-for-reproducibility
Markdown
/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-data-version-control-for-reproducibility.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-data-version-control-for-reproducibility/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-data-version-control-for-reproducibility.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

Lucas and Luna break down why data version control (DVC) has become as essential as Git for machine learning teams. They trace the problem through a concrete example: a fraud detection model at a fintech company where a missing dataset version caused a 15 percent drop in recall. The episode walks through how DVC tracks data snapshots, pipeline stages, and model artifacts—without duplicating massive files—using a simple declarative YAML config. Lucas explains the difference between DVC's approach and Git LFS, and why tools like Pachyderm and DVC solve overlapping but distinct problems. The hosts also discuss how versioning interacts with feature stores and CI/CD for ML, and why the field is moving toward treating data with the same discipline as source code. No fluff, just a focused look at one practice that separates professional data teams from the rest. #DataVersionControl #DVC #MLOps #Reproducibility #MachineLearning #DataScience #GitForData #Pachyderm #LFS #DataPipeline #FeatureStore #CI/CD #FraudDetection #Fintech #MLPipeline #DataGovernance #Technology #FexingoBusiness Keep every episode free: buymeacoffee.com/fexingo