{"podcast":{"title":"The Data Science Podcast with Fexingo: Analytics, Machine Learning, and Data-Driven Conversations","slug":"the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831","podcast_index_feed_id":7871831,"rss_url":"https://feeds.fexingo.com/business/the-data-science-podcast.xml","website_url":"https://www.fexingo.com/","image_url":"https://audio.fexingo.com/business/the-data-science-podcast/cover.png","author":"Fexingo","episode_count":118,"summary":"Lucas and Luna sit at a data-science workstation, two thin laptops open to scatter plots and clustering visualizations, and ask: what can we actually learn from the numbers? Each episode of The Data Science Podcast with Fexingo is a grounded, specific conversation about a single analytics problem or machine-learning method — from regularization in regression to the bias-variance trade-off in random forests. Lucas leads with a journalistic eye for how models are built and tested in the real world, citing actual case studies like how Netflix used matrix factorization for recommendations or how healthcare researchers apply survival analysis to clinical trials. Luna keeps the discussion honest, asking about data quality, feature engineering pitfalls, and whether a model’s accuracy actually translates to business value. They never resort to buzzwords: instead, they walk through the workflow from data collection to deployment, discussing trade-offs like interpretability versus performance. The show serves data scientists, analysts, and engineers who want to stay sharp on methods without the hype. Listeners walk away with a clearer understanding of why one algorithm beats another on a gi…","last_synced_at":"2026-07-19T08:17:23.323447+00:00","page_url":"https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831"},"episode":{"title":"How Data Scientists Use Synthetic Data for Model Training","slug":"how-data-scientists-use-synthetic-data-for-model-training-2","published_at":"2026-07-09T21:03:26+00:00","page_url":"https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-synthetic-data-for-model-training-2","show_page_url":"https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831","url":"https://audio.fexingo.com/business/the-data-science-podcast/episode-0100.mp3","audio_url":"https://audio.fexingo.com/business/the-data-science-podcast/episode-0100.mp3","summary":"For Episode 100 of The Data Science Podcast, Lucas and Luna explore the controversial practice of training AI models on synthetic data. With real-world examples from autonomous vehicle companies like Waymo and medical imaging startups, they discuss when synthetic data works, when it fails, and how the 'data cascade' problem threatens to pollute next-generation models. Lucas explains why some researchers call synthetic data a 'tax' on human-generated data, and Luna pushes back on whether it can ever truly replace the real thing. A grounded, numbers-driven conversation for anyone building or buying data products in 2026. #SyntheticData #DataScience #MachineLearning #AI #Waymo #AutonomousVehicles #MedicalImaging #DataAugmentation #GANs #DiffusionModels #ModelCollapse #DataCascade #TrainingData #Tech #FexingoBusiness #BusinessPodcast #Episode100 #GenerativeAI Keep every episode free: buymeacoffee.com/fexingo","meta_description":"For Episode 100 of The Data Science Podcast, Lucas and Luna explore the controversial practice of training AI models on synthetic data. With real-world ex…","key_points":[],"chapters":[],"topics":[],"duration_seconds":491,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-synthetic-data-for-model-training-2/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-synthetic-data-for-model-training-2.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}