# MLOps Week 18: The LLM Revolution & the Future of Data with Josh Wills Page: https://stenobird.com/podcast/mlops-weekly/mlops-week-18-the-llm-revolution-the-future-of-data-with-josh-wills Text version: https://stenobird.com/podcast/mlops-weekly/mlops-week-18-the-llm-revolution-the-future-of-data-with-josh-wills.md Podcast: [MLOps Weekly Podcast](https://stenobird.com/podcast/mlops-weekly) Published: 2023-05-24T15:37:13+00:00 Episode link: https://rss.com/podcasts/mlops-weekly/966178 Audio file: https://content.rss.com/episodes/132586/966178/mlops-weekly/2023_05_24_15_36_48_de6a6a90-4335-4dff-a632-ca6d985881f9.mp3 Processing state: processed JSON: https://stenobird.com/v1/public/podcasts/mlops-weekly/episodes/mlops-week-18-the-llm-revolution-the-future-of-data-with-josh-wills Duration seconds: 2773 ## Resource Josh Wills shares lessons from scaling data pipelines at Slack and Google, focusing on the necessity of data contracts. The discussion explores whether LLMs represent a fundamental shift like electricity or a functional shift like the mobile phone. ## Highlights - Main idea: Data contracts are essential for decoupling production systems from downstream data warehouses - Practical takeaway: Treat data pipelines as production-grade software by implementing integration tests between upstream and downstream systems - Failure mode: Over-identifying with source code can lead to professional burnout and resistance to necessary code reviews - Main idea: The future of LLMs likely depends on whether the economics favor a single dominant model or a fragmented ecosystem of specialized models - Practical takeaway: In high-scale environments, prepare for 'one-in-a-trillion' edge cases that become frequent when processing trillions of records ## Topics Data Engineering, LLMs, Data Contracts, MLOps, Scalable Systems, Data Warehousing, Artificial Intelligence, GPU Scarcity ## Chapters - 1:00 — The Mount Everest of Data Engineering: Reflecting on the technical and human challenges of rebuilding Slack's massive search indexing pipeline. - 4:20 — Implementing Data Contracts: How using Thrift schemas to tie production systems to the data warehouse prevents downstream breakage. - 18:20 — The Hype Cycle of LLMs: Evaluating whether Large Language Models are a paradigm shift akin to the mobile phone or the invention of electricity. - 21:45 — Lessons from Google and Go: A look at the adoption of the Go programming language and the importance of tools fitting the community's needs. - 25:25 — The Middle Ground of AI Hype: Navigating the space between extreme skepticism and the 'electricity-level' hype of generative AI. - 35:40 — The Economics of Model Training: Discussing the scarcity of GPUs and the business incentives for platforms like Databricks to support decentralized model training. ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/mlops-weekly/episodes/mlops-week-18-the-llm-revolution-the-future-of-data-with-josh-wills/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/mlops-weekly/mlops-week-18-the-llm-revolution-the-future-of-data-with-josh-wills.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.