Episode

Digital Dialogs (Episode 3 | S5) - Foundation Models for Computer Vision

Podcast
Digital Dialogs - Compact conversations on the future of tech
Published
Jun 4, 2026
Duration seconds
1497
Processing state
not_requested
Canonical source
https://www.reply.com/en/artificial-intelligence/digital-dialogues
Audio
https://episodes.captivate.fm/episode/0e88bd0a-8d19-4d81-bb13-bdd30cad5bf2.mp3
JSON
/v1/public/podcasts/digital-dialogs-compact-conversations-on-the-future-of-tech-7066827/episodes/digital-dialogs-episode-3-s5-foundation-models-for-computer-vision
Markdown
/podcast/digital-dialogs-compact-conversations-on-the-future-of-tech-7066827/digital-dialogs-episode-3-s5-foundation-models-for-computer-vision.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/digital-dialogs-compact-conversations-on-the-future-of-tech-7066827/episodes/digital-dialogs-episode-3-s5-foundation-models-for-computer-vision/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/digital-dialogs-compact-conversations-on-the-future-of-tech-7066827/digital-dialogs-episode-3-s5-foundation-models-for-computer-vision.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

In this episode, Johannes Morzeck explores how Vision Foundation Models (VLMs) and specialized industrial models are fundamentally transforming computer vision from expensive custom R&D projects into scalable, ready-to-use business solutions. Johannes discusses the real breakthroughs compared to classic computer vision and the practical challenges organizations face when moving these massive models from proof-of-concept to production in real-world environments like factories and retail stores. The conversation delves into the complexities of balancing model performance with the cost of running them at the edge. Looking toward the future, Johannes shares his insights on how these models will reshape industries reliant on visual data, envisioning a world where cameras serve as "universal sensors" capable of answering any question about physical operations without the need for custom training.