Episode

The Professor of Outputmaxxing — Anjney Midha, AMP

Podcast
Latent Space: The AI Engineer Podcast
Published
Jun 18, 2026
Duration seconds
3565
Processing state
processed
Canonical source
https://www.latent.space/p/anj
Audio
https://api.substack.com/feed/podcast/202359797/7d6863592b786561d5ce8ed820585ddb.mp3
JSON
/v1/public/podcasts/latent-space-ai-engineer/episodes/the-professor-of-outputmaxxing-anjney-midha-amp
Markdown
/podcast/latent-space-ai-engineer/the-professor-of-outputmaxxing-anjney-midha-amp.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/latent-space-ai-engineer/episodes/the-professor-of-outputmaxxing-anjney-midha-amp/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/latent-space-ai-engineer/the-professor-of-outputmaxxing-anjney-midha-amp.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

The AI scaling race is often framed as a pursuit of more GPUs, but the real frontier lies in maximizing Model FLOPs Utilization (MFU). Anjney Midha argues that inefficient cluster management and misaligned incentives are causing massive computational waste.

Topics

  • AI Infrastructure
  • Model FLOPs Utilization
  • GPU Scaling
  • Compute Grids
  • Systems Engineering
  • Decentralized Computing
  • Machine Learning Operations
  • Hardware-Software Co-design

Highlights

  • Main idea: Scaling AI is increasingly a systems engineering problem involving scheduling, networking, and kernels rather than just raw CapEx
  • Failure mode: Low Model FLOPs Utilization (MFU) in frontier labs suggests that simply adding more GPUs won't yield proportional progress without better orchestration
  • Practical takeaway: The future of compute lies in a decentralized, protocol-based grid where supply and demand can flow like a utility
  • Main idea: High-performance clusters should aim for much higher utilization rates, noting that 95% node utilization is a standard for reliability at scale
  • Strategic insight: To enable new hardware, developers should adopt the NVIDIA reference architecture to ensure compatibility with existing software stacks

Chapters

  1. 1:00 The Alignment of Compute and Capital: An analysis of how the gap between funding and deployment leads to massive inefficiencies and wasted compute in large-scale clusters.
  2. 10:00 Building In-House Infrastructure: Lessons from Discord's approach to building proprietary communication infrastructure to avoid third-party bottlenecks.
  3. 23:00 AI for End-of-Life Prediction: Exploring how AI can be applied to clinical decision-making and reducing the societal burden of healthcare through science.
  4. 32:00 The Importance of Co-design: Why hardware and software developers need visibility into future model generations to prevent architectural mismatches.
  5. 37:00 The AMP Compute Grid Vision: How AMP uses excess compute to support non-profits and universities while building a scalable infrastructure model.
  6. 50:00 Mission Alignment and Culture: A discussion on maintaining organizational integrity and mission-driven values in the face of rapid scaling and mercenary incentives.