Episode

MLOps Week 18: The LLM Revolution & the Future of Data with Josh Wills

Podcast
MLOps Weekly Podcast
Published
May 24, 2023
Duration seconds
2773
Processing state
processed
Canonical source
https://rss.com/podcasts/mlops-weekly/966178
Audio
https://content.rss.com/episodes/132586/966178/mlops-weekly/2023_05_24_15_36_48_de6a6a90-4335-4dff-a632-ca6d985881f9.mp3
JSON
/v1/public/podcasts/mlops-weekly/episodes/mlops-week-18-the-llm-revolution-the-future-of-data-with-josh-wills
Markdown
/podcast/mlops-weekly/mlops-week-18-the-llm-revolution-the-future-of-data-with-josh-wills.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/mlops-weekly/episodes/mlops-week-18-the-llm-revolution-the-future-of-data-with-josh-wills/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/mlops-weekly/mlops-week-18-the-llm-revolution-the-future-of-data-with-josh-wills.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

Josh Wills shares lessons from scaling data pipelines at Slack and Google, focusing on the necessity of data contracts. The discussion explores whether LLMs represent a fundamental shift like electricity or a functional shift like the mobile phone.

Topics

  • Data Engineering
  • LLMs
  • Data Contracts
  • MLOps
  • Scalable Systems
  • Data Warehousing
  • Artificial Intelligence
  • GPU Scarcity

Highlights

  • Main idea: Data contracts are essential for decoupling production systems from downstream data warehouses
  • Practical takeaway: Treat data pipelines as production-grade software by implementing integration tests between upstream and downstream systems
  • Failure mode: Over-identifying with source code can lead to professional burnout and resistance to necessary code reviews
  • Main idea: The future of LLMs likely depends on whether the economics favor a single dominant model or a fragmented ecosystem of specialized models
  • Practical takeaway: In high-scale environments, prepare for 'one-in-a-trillion' edge cases that become frequent when processing trillions of records

Chapters

  1. 1:00 The Mount Everest of Data Engineering: Reflecting on the technical and human challenges of rebuilding Slack's massive search indexing pipeline.
  2. 4:20 Implementing Data Contracts: How using Thrift schemas to tie production systems to the data warehouse prevents downstream breakage.
  3. 18:20 The Hype Cycle of LLMs: Evaluating whether Large Language Models are a paradigm shift akin to the mobile phone or the invention of electricity.
  4. 21:45 Lessons from Google and Go: A look at the adoption of the Go programming language and the importance of tools fitting the community's needs.
  5. 25:25 The Middle Ground of AI Hype: Navigating the space between extreme skepticism and the 'electricity-level' hype of generative AI.
  6. 35:40 The Economics of Model Training: Discussing the scarcity of GPUs and the business incentives for platforms like Databricks to support decentralized model training.