Episode
How Data Scientists Use Synthetic Data for Model Training
- Podcast
- The Data Science Podcast with Fexingo: Analytics, Machine Learning, and Data-Driven Conversations
- Published
- Jun 18, 2026
- Duration seconds
- 619
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-synthetic-data-for-model-training/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-synthetic-data-for-model-training.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
Episode 58 of The Data Science Podcast explores how data scientists are turning to synthetic data when real-world data is scarce, expensive, or privacy-sensitive. Lucas and Luna break down a concrete case: how a European insurance company used the Synthpop library in R to generate synthetic claims data, training a fraud detection model that outperformed their original one. They discuss the core techniques—generative models like GANs and VAEs, as well as statistical synthetization methods—and the trade-offs around fidelity, privacy, and bias. The episode also covers the rise of synthetic data platforms like Mostly AI and Gretel.ai, and why synthetic data is becoming a standard tool for teams that need to augment small datasets or test edge cases without exposing real user information. No hype, just what works and what doesn't. #SyntheticData #MachineLearning #DataScience #GANs #VAEs #Synthpop #MostlyAI #GretelAI #FraudDetection #InsuranceAnalytics #PrivacyPreservingML #DataAugmentation #GenerativeModels #SyntheticDataPlatforms #Technology #FexingoBusiness #BusinessPodcast #DataDriven Keep every episode free: buymeacoffee.com/fexingo