# How Data Scientists Use Multimodal Models for Zero-Shot Learning Page: https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-multimodal-models-for-zero-shot-learning Text version: https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-multimodal-models-for-zero-shot-learning.md Podcast: [The Data Science Podcast with Fexingo: Analytics, Machine Learning, and Data-Driven Conversations](https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831) Published: 2026-07-07T21:00:43+00:00 Episode link: https://audio.fexingo.com/business/the-data-science-podcast/episode-0096.mp3 Audio file: https://audio.fexingo.com/business/the-data-science-podcast/episode-0096.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-multimodal-models-for-zero-shot-learning Duration seconds: 704 ## Resource In this episode, Lucas and Luna dive into multimodal zero-shot learning, the technique that lets AI models like CLIP recognize objects, scenes, and text across images without ever being explicitly trained on those combinations. They explore a concrete use case: a retail startup using a pretrained multimodal embedding model to automatically tag 10,000 new product photos per day with zero labeling cost. Lucas breaks down the architecture—contrastive learning on image-text pairs—and explains why zero-shot works when the embedding space is aligned. Luna asks about failure modes: ambiguous images, domain shift, adversarial inputs. They also touch on the trade-off between generality and fine-tuning, and where the field is heading next. No hype, just how the math makes it possible. #MultimodalLearning #ZeroShotLearning #CLIP #ContrastiveLearning #DataScience #MachineLearning #ImageRecognition #NaturalLanguageProcessing #Embeddings #AI #RetailTech #ProductTagging #TransferLearning #DeepLearning #ComputerVision #Technology #FexingoBusiness #BusinessPodcast Keep every episode free: buymeacoffee.com/fexingo ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-multimodal-models-for-zero-shot-learning/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-multimodal-models-for-zero-shot-learning.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.