Episode
Data Scientists Use Active Learning to Label Smarter
- Podcast
- The Data Science Podcast with Fexingo: Analytics, Machine Learning, and Data-Driven Conversations
- Published
- Jul 6, 2026
- Duration seconds
- 553
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/data-scientists-use-active-learning-to-label-smarter/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/data-scientists-use-active-learning-to-label-smarter.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
In episode 94 of The Data Science Podcast with Fexingo, Lucas and Luna explore how active learning cuts labeling costs by 80 percent while maintaining model accuracy. Using a concrete example from a medical imaging startup training a rare-disease classifier, they walk through uncertainty sampling, query strategies, and the human-in-the-loop workflow. They compare pool-based versus stream-based active learning, discuss common pitfalls like distribution shift, and explain when active learning beats random sampling. If you are a data scientist looking to stretch a limited labeling budget, this episode gives you a practical framework to get started. Lucas and Luna also touch on tools like modAL, scikit-activeml, and Label Studio. No hype, just signal. #ActiveLearning #DataScience #MachineLearning #Labeling #UncertaintySampling #HumanInTheLoop #MedicalImaging #RareDisease #modAL #scikitActivelm #LabelStudio #DistributionShift #QueryStrategy #PoolBased #StreamBased #Technology #FexingoBusiness #BusinessPodcast Keep every episode free: buymeacoffee.com/fexingo