Episode

How Data Scientists Use Multimodal Models for Zero-Shot Learning

Podcast
The Data Science Podcast with Fexingo: Analytics, Machine Learning, and Data-Driven Conversations
Published
Jul 7, 2026
Duration seconds
704
Processing state
not_requested
Canonical source
https://audio.fexingo.com/business/the-data-science-podcast/episode-0096.mp3
Audio
https://audio.fexingo.com/business/the-data-science-podcast/episode-0096.mp3
JSON
/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-multimodal-models-for-zero-shot-learning
Markdown
/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-multimodal-models-for-zero-shot-learning.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-multimodal-models-for-zero-shot-learning/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-multimodal-models-for-zero-shot-learning.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

In this episode, Lucas and Luna dive into multimodal zero-shot learning, the technique that lets AI models like CLIP recognize objects, scenes, and text across images without ever being explicitly trained on those combinations. They explore a concrete use case: a retail startup using a pretrained multimodal embedding model to automatically tag 10,000 new product photos per day with zero labeling cost. Lucas breaks down the architecture—contrastive learning on image-text pairs—and explains why zero-shot works when the embedding space is aligned. Luna asks about failure modes: ambiguous images, domain shift, adversarial inputs. They also touch on the trade-off between generality and fine-tuning, and where the field is heading next. No hype, just how the math makes it possible. #MultimodalLearning #ZeroShotLearning #CLIP #ContrastiveLearning #DataScience #MachineLearning #ImageRecognition #NaturalLanguageProcessing #Embeddings #AI #RetailTech #ProductTagging #TransferLearning #DeepLearning #ComputerVision #Technology #FexingoBusiness #BusinessPodcast Keep every episode free: buymeacoffee.com/fexingo