Episode

Inside LinkedIn’s AI Engineering Playbook

Podcast
Beyond The Pilot: Enterprise AI in Action
Published
Jan 21, 2026
Duration seconds
2438
Processing state
not_requested
Canonical source
https://traffic.megaphone.fm/UTEAU9861507389.mp3?updated=1768929450
Audio
https://traffic.megaphone.fm/UTEAU9861507389.mp3?updated=1768929450
JSON
/v1/public/podcasts/beyond-the-pilot-enterprise-ai-in-action-7531113/episodes/inside-linkedin-s-ai-engineering-playbook
Markdown
/podcast/beyond-the-pilot-enterprise-ai-in-action-7531113/inside-linkedin-s-ai-engineering-playbook.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/beyond-the-pilot-enterprise-ai-in-action-7531113/episodes/inside-linkedin-s-ai-engineering-playbook/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/beyond-the-pilot-enterprise-ai-in-action-7531113/inside-linkedin-s-ai-engineering-playbook.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

While the rest of the industry chases massive models, LinkedIn quietly achieved a major engineering breakthrough by going small. In this episode of Beyond the Pilot, Erran Berger (VP of Product Engineering, LinkedIn) opens the "cookbook" on how they distilled massive 7B parameter models down to ultra-efficient 600M parameter "student" models—scaling AI to 1.2 billion users without breaking the bank. AI Gets Real Here. This isn't theory. Erran details the exact architecture, the "Multi-Teacher" distillation process, and the organizational shift that forced Product Managers to write evals instead of specs. In this episode, we cover: The Distillation Pipeline: How to train a 7B "Teacher" and distill it to a 1.7B intermediate and 0.6B "Student" for production. Synthetic Data Strategy: Using GPT-4 to generate the "Golden Dataset" for training. Multi-Teacher Architecture: Why they separated "Product Policy" and "Click Prediction" into different teacher models to solve alignment issues. 10x Efficiency Hacks: Specific techniques (Pruning, Quantization, Context Compression) that slashed latency. Org Design: Why the "Eval First" culture is the new requirement for AI engineering teams. 🚀 CHAPTERS 00:00 - Intro: LinkedIn's Massive "Small Model" Feat 04:00 - Why Commercial Models Failed at LinkedIn Scale 08:00 - The "Product Policy" Funnel & Synthetic Data Generation 12:00 - The Pipeline: 7B → 1.7B → 600M Parameters 19:00 - The "Multi-Teacher" Breakthrough (Relevance vs. Clicks) 23:00 - How They Achieved 10x Latency Reduction (Pruning/Compression) 31:00 - Changing the Culture: Why PMs Must Write Evals 35:00 - The "Bright Green Matrix": Measuring Success & Future Roadmap Presented by Outshift by Cisco Outshift is Cisco’s emerging tech incubation engine and driver of Agentic AI, quan…