Episode
How Data Scientists Use Retrieval Augmented Generation for Enterprise Search
- Podcast
- The Data Science Podcast with Fexingo: Analytics, Machine Learning, and Data-Driven Conversations
- Published
- Jul 14, 2026
- Duration seconds
- 472
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-retrieval-augmented-generation-for-enterprise-search/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-retrieval-augmented-generation-for-enterprise-search.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
Episode 110 of The Data Science Podcast dives into Retrieval Augmented Generation (RAG) for enterprise search. Lucas and Luna explore how companies like JP Morgan and NASA are using RAG to make internal documents searchable and actionable. They discuss the key components: embedding models, vector databases like Pinecone, and large language models like GPT-4. The episode walks through a concrete example: a financial analyst querying a 10-K filing for revenue recognition policies. They cover challenges like chunking strategies, retrieval quality, and hallucination risks, plus emerging techniques like HyDE and multi-hop retrieval. By the end, listeners understand RAG's role in unlocking unstructured data at scale. #RetrievalAugmentedGeneration #EnterpriseSearch #RAG #VectorDatabases #Embeddings #LargeLanguageModels #GPT4 #Pinecone #JP Morgan #NASA #10K Filing #HyDE #MultiHopRetrieval #UnstructuredData #DataScience #Technology #FexingoBusiness #BusinessPodcast Keep every episode free: buymeacoffee.com/fexingo