{"podcast":{"title":"Beyond The Pilot: Enterprise AI in Action","slug":"beyond-the-pilot-enterprise-ai-in-action-7531113","podcast_index_feed_id":7531113,"rss_url":"https://feeds.megaphone.fm/UTEAU6270845142","website_url":null,"image_url":"https://megaphone.imgix.net/podcasts/2eacf1e6-89b2-11f0-8a6f-13263c2b3be9/image/303aa3d4fcd77adce1002a0f6735a3fb.png?ixlib=rails-4.3.1&max-w=3000&max-h=3000&fit=crop&auto=format,compress","author":"VentureBeat","episode_count":32,"summary":"AI gets real here. On “Beyond the Pilot,” top business execs share what actually happens after the AI proof of concept — from infrastructure and org design to wins, failures, and ROI. Not theory, but deep dives into how they scaled AI that works.","last_synced_at":"2026-06-24T22:19:17.509819+00:00","page_url":"https://stenobird.com/podcast/beyond-the-pilot-enterprise-ai-in-action-7531113"},"episode":{"title":"GPU Hoarding is Over. The $401B Reality Check","slug":"gpu-hoarding-is-over-the-401b-reality-check","published_at":"2026-05-13T10:29:00+00:00","page_url":"https://stenobird.com/podcast/beyond-the-pilot-enterprise-ai-in-action-7531113/gpu-hoarding-is-over-the-401b-reality-check","show_page_url":"https://stenobird.com/podcast/beyond-the-pilot-enterprise-ai-in-action-7531113","url":"https://traffic.megaphone.fm/UTEAU2517377367.mp3","audio_url":"https://traffic.megaphone.fm/UTEAU2517377367.mp3","summary":"Enterprise GPU hoarding is over. LinkedIn CTO Erran Berger and VentureBeat analyst Rob Strechay break down what comes next — and the infrastructure math most enterprises are only now being forced to confront. VentureBeat's Q1 research shows GPU availability anxiety dropped from 20.8% to 15.4% among enterprise teams, while cost-per-inference and TCO concerns jumped from 34% to 41% — a number that's still climbing. The hoarding phase is giving way to an audit phase, and the companies that didn't build the instrumentation to understand their workloads are now paying for it. Erran Berger explains how LinkedIn runs one of the few remaining at-scale applied ML shops outside the hyperscalers — owning the full stack from bare metal GPU clusters to member-facing products. That means LinkedIn engineers can optimize custom CUDA kernels, compress embeddings, prune models for throughput, and adapt networking and storage per workload — trade-offs that are simply unavailable on public cloud instance menus. The result: a rigorous ROI framework that evaluates not just current traffic costs, but the traffic shape agents will drive in 2–3 years. On the market side, 72% of enterprises admit they lack sufficient control over their AI infrastructure. Open-source inference tools like vLLM and LLMD are seeing rapid adoption, while 17% of organizations have moved to full-stack ownership. Hyperscalers report 60–80% of workloads have already shifted from training to inference — and most enterprise teams are still figuring out how to staff and instrument for that reality. 🎙️ GUEST: Erran Berger | CTO, LinkedIn 🎙️ ANALYST: Rob Strechay | VentureBeat 🎙️ HOST: Matt Marshall | CEO, VentureBeat --- 00:00 Intro: The GPU Hoarding Hangover 00:10 Guest Introductions 02:00 VentureBeat Q1 Data: GPU Panic Fa…","meta_description":"Enterprise GPU hoarding is over. LinkedIn CTO Erran Berger and VentureBeat analyst Rob Strechay break down what comes next — and the infrastructure math m…","key_points":[],"chapters":[],"topics":[],"duration_seconds":958,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/beyond-the-pilot-enterprise-ai-in-action-7531113/episodes/gpu-hoarding-is-over-the-401b-reality-check/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/beyond-the-pilot-enterprise-ai-in-action-7531113/gpu-hoarding-is-over-the-401b-reality-check.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}