# How Data Scientists Use Synthetic Data for Model Training Page: https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-synthetic-data-for-model-training Text version: https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-synthetic-data-for-model-training.md Podcast: [The Data Science Podcast with Fexingo: Analytics, Machine Learning, and Data-Driven Conversations](https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831) Published: 2026-06-18T08:15:43+00:00 Episode link: https://audio.fexingo.com/business/the-data-science-podcast/episode-0058.mp3 Audio file: https://audio.fexingo.com/business/the-data-science-podcast/episode-0058.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-synthetic-data-for-model-training Duration seconds: 619 ## Resource Episode 58 of The Data Science Podcast explores how data scientists are turning to synthetic data when real-world data is scarce, expensive, or privacy-sensitive. Lucas and Luna break down a concrete case: how a European insurance company used the Synthpop library in R to generate synthetic claims data, training a fraud detection model that outperformed their original one. They discuss the core techniques—generative models like GANs and VAEs, as well as statistical synthetization methods—and the trade-offs around fidelity, privacy, and bias. The episode also covers the rise of synthetic data platforms like Mostly AI and Gretel.ai, and why synthetic data is becoming a standard tool for teams that need to augment small datasets or test edge cases without exposing real user information. No hype, just what works and what doesn't. #SyntheticData #MachineLearning #DataScience #GANs #VAEs #Synthpop #MostlyAI #GretelAI #FraudDetection #InsuranceAnalytics #PrivacyPreservingML #DataAugmentation #GenerativeModels #SyntheticDataPlatforms #Technology #FexingoBusiness #BusinessPodcast #DataDriven Keep every episode free: buymeacoffee.com/fexingo ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-synthetic-data-for-model-training/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-synthetic-data-for-model-training.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.