Episode
How Data Scientists Use Data Version Control for Reproducibility
- Podcast
- The Data Science Podcast with Fexingo: Analytics, Machine Learning, and Data-Driven Conversations
- Published
- Jul 13, 2026
- Duration seconds
- 756
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-data-version-control-for-reproducibility/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-data-version-control-for-reproducibility.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
Lucas and Luna break down why data version control (DVC) has become as essential as Git for machine learning teams. They trace the problem through a concrete example: a fraud detection model at a fintech company where a missing dataset version caused a 15 percent drop in recall. The episode walks through how DVC tracks data snapshots, pipeline stages, and model artifacts—without duplicating massive files—using a simple declarative YAML config. Lucas explains the difference between DVC's approach and Git LFS, and why tools like Pachyderm and DVC solve overlapping but distinct problems. The hosts also discuss how versioning interacts with feature stores and CI/CD for ML, and why the field is moving toward treating data with the same discipline as source code. No fluff, just a focused look at one practice that separates professional data teams from the rest. #DataVersionControl #DVC #MLOps #Reproducibility #MachineLearning #DataScience #GitForData #Pachyderm #LFS #DataPipeline #FeatureStore #CI/CD #FraudDetection #Fintech #MLPipeline #DataGovernance #Technology #FexingoBusiness Keep every episode free: buymeacoffee.com/fexingo