# How Data Scientists Use Vector Databases for RAG Systems Page: https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-vector-databases-for-rag-systems Text version: https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-vector-databases-for-rag-systems.md Podcast: [The Data Science Podcast with Fexingo: Analytics, Machine Learning, and Data-Driven Conversations](https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831) Published: 2026-06-19T20:45:11+00:00 Episode link: https://audio.fexingo.com/business/the-data-science-podcast/episode-0061.mp3 Audio file: https://audio.fexingo.com/business/the-data-science-podcast/episode-0061.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-vector-databases-for-rag-systems Duration seconds: 676 ## Resource Retrieval-augmented generation, or RAG, is reshaping how companies deploy large language models without retraining. In this episode, Lucas and Luna drill into the data-science architecture behind RAG: how vector databases encode semantic meaning, why cosine similarity beats keyword search, and what a production RAG pipeline looks like at a mid-size fintech startup. They walk through a concrete example—building a customer-support bot for a payments company—showing where embedding models, chunking strategies, and approximate nearest-neighbor search come into play. Lucas breaks down the trade-off between accuracy and latency, and Luna questions whether RAG is just a band-aid for models that can't reason. Tune in for a grounded look at the database layer that's quietly powering the next wave of AI applications. #VectorDatabases #RAG #RetrievalAugmentedGeneration #LLM #Embeddings #CosineSimilarity #ApproximateNearestNeighbor #DataScience #MachineLearning #NLP #AIArchitecture #CustomerSupport #Fintech #Pinecone #Weaviate #Milvus #FexingoBusiness #TechnologyPodcast Keep every episode free: buymeacoffee.com/fexingo ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-vector-databases-for-rag-systems/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-vector-databases-for-rag-systems.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.