Episode
How Data Scientists Use Synthetic Data for Model Training
- Podcast
- The Data Science Podcast with Fexingo: Analytics, Machine Learning, and Data-Driven Conversations
- Published
- Jul 9, 2026
- Duration seconds
- 491
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-synthetic-data-for-model-training-2/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-synthetic-data-for-model-training-2.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
For Episode 100 of The Data Science Podcast, Lucas and Luna explore the controversial practice of training AI models on synthetic data. With real-world examples from autonomous vehicle companies like Waymo and medical imaging startups, they discuss when synthetic data works, when it fails, and how the 'data cascade' problem threatens to pollute next-generation models. Lucas explains why some researchers call synthetic data a 'tax' on human-generated data, and Luna pushes back on whether it can ever truly replace the real thing. A grounded, numbers-driven conversation for anyone building or buying data products in 2026. #SyntheticData #DataScience #MachineLearning #AI #Waymo #AutonomousVehicles #MedicalImaging #DataAugmentation #GANs #DiffusionModels #ModelCollapse #DataCascade #TrainingData #Tech #FexingoBusiness #BusinessPodcast #Episode100 #GenerativeAI Keep every episode free: buymeacoffee.com/fexingo