# Drilling Down on Data with Bobby Neelon & John Kalfayan (Collide) Page: https://stenobird.com/podcast/chuck-yates-got-a-job-1369602/drilling-down-on-data-with-bobby-neelon-john-kalfayan-collide Text version: https://stenobird.com/podcast/chuck-yates-got-a-job-1369602/drilling-down-on-data-with-bobby-neelon-john-kalfayan-collide.md Podcast: [Chuck Yates Got A Job](https://stenobird.com/podcast/chuck-yates-got-a-job-1369602) Published: 2025-11-25T13:00:00+00:00 Episode link: https://cynaj.collide.io/episodes/drilling-down-on-data-with-bobby-neelon-john-kalfayan-collide Audio file: https://media.transistor.fm/88e58198/f056a02f.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/chuck-yates-got-a-job-1369602/episodes/drilling-down-on-data-with-bobby-neelon-john-kalfayan-collide Duration seconds: 2486 ## Resource Bobby Neelon and John Kalfayan from Collide break down the messy reality of getting data ready for RAG, why PDFs are dumpster fires for unstructured data, how extraction changes depending on whether you're dealing with drilling surveys or handwritten logs, and why chunking strategy matters more than people think. They walk through embeddings, vector databases, MCP servers for pulling external data without leaking internal info, and why good metadata and folder structure actually make AI deployments way easier. Plus the hard truth that AI isn't a silver bullet for bad data management and the crap-in-crap-out problem is getting worse because now it can hallucinate on top of the crap. Click here to watch a video of this episode. Join the conversation shaping the future of energy. Collide is the community where oil & gas professionals connect, share insights, and solve real-world problems together. No noise. No fluff. Just the discussions that move our industry forward. Apply today at collide.io Click here to view the episode transcript. 0:00 - Introductions and RAG overview 3:15 - Document identification and classification challenges 8:40 - Extracting data from unstructured PDFs 13:25 - Real world examples of messy data formats 18:50 - OCR paired with vision models for extraction 22:10 - Chunking strategies and when to use each 26:35 - Embeddings and vector databases explained 30:20 - MCP servers and external data integration 35:45 - Getting data AI-ready with metadata and structure 40:30 - Text-to-SQL approaches and database access 44:15 - Handling duplicates and M&A data integration 48:50 - How AI learns context over time 53:40 - Why traditional data management matters more than ever https://twitter.com/collide_io https://www.tiktok.com/@collide.io https://www.f… ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/chuck-yates-got-a-job-1369602/episodes/drilling-down-on-data-with-bobby-neelon-john-kalfayan-collide/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/chuck-yates-got-a-job-1369602/drilling-down-on-data-with-bobby-neelon-john-kalfayan-collide.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.