# Why AI Model Inference Is Splitting Into Two Markets Page: https://stenobird.com/podcast/the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011/why-ai-model-inference-is-splitting-into-two-markets Text version: https://stenobird.com/podcast/the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011/why-ai-model-inference-is-splitting-into-two-markets.md Podcast: [The AI Podcast with Fexingo: Artificial Intelligence, Machine Learning, and Modern AI Models](https://stenobird.com/podcast/the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011) Published: 2026-07-06T20:28:22+00:00 Episode link: https://audio.fexingo.com/business/the-ai-podcast/episode-0095.mp3 Audio file: https://audio.fexingo.com/business/the-ai-podcast/episode-0095.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011/episodes/why-ai-model-inference-is-splitting-into-two-markets Duration seconds: 551 ## Resource Lucas and Luna explore the surprising divergence in AI infrastructure as model inference splits into two distinct markets: high-throughput batch processing for enterprise and low-latency real-time inference for consumer apps. They discuss how companies like Together AI and Anthropic are betting on custom silicon, why NVIDIA's data center revenue mix is shifting, and what the recent sell-off in equipment makers like ASML and Applied Materials signals for the next wave of AI deployment. The hosts also unpack a recent interview with Vercel's CEO on the fight to separate models from agents, and what it means for startups building on top of foundation models. #AIInference #ModelSplitting #NVIDIA #TogetherAI #Anthropic #Vercel #ASML #AppliedMaterials #CustomSilicon #AIInfrastructure #LatencyVsThroughput #RealTimeAI #BatchProcessing #AIStartups #Technology #FexingoBusiness #BusinessPodcast #AIPodcast Keep every episode free: buymeacoffee.com/fexingo ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011/episodes/why-ai-model-inference-is-splitting-into-two-markets/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011/why-ai-model-inference-is-splitting-into-two-markets.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.