Episode
Why AI Model Inference Is Splitting Into Two Markets
- Published
- Jul 6, 2026
- Duration seconds
- 551
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011/episodes/why-ai-model-inference-is-splitting-into-two-markets/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011/why-ai-model-inference-is-splitting-into-two-markets.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
Lucas and Luna explore the surprising divergence in AI infrastructure as model inference splits into two distinct markets: high-throughput batch processing for enterprise and low-latency real-time inference for consumer apps. They discuss how companies like Together AI and Anthropic are betting on custom silicon, why NVIDIA's data center revenue mix is shifting, and what the recent sell-off in equipment makers like ASML and Applied Materials signals for the next wave of AI deployment. The hosts also unpack a recent interview with Vercel's CEO on the fight to separate models from agents, and what it means for startups building on top of foundation models. #AIInference #ModelSplitting #NVIDIA #TogetherAI #Anthropic #Vercel #ASML #AppliedMaterials #CustomSilicon #AIInfrastructure #LatencyVsThroughput #RealTimeAI #BatchProcessing #AIStartups #Technology #FexingoBusiness #BusinessPodcast #AIPodcast Keep every episode free: buymeacoffee.com/fexingo