{"podcast":{"title":"The AI Podcast with Fexingo: Artificial Intelligence, Machine Learning, and Modern AI Models","slug":"the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011","podcast_index_feed_id":7872011,"rss_url":"https://feeds.fexingo.com/business/the-ai-podcast.xml","website_url":"https://www.fexingo.com/","image_url":"https://audio.fexingo.com/business/the-ai-podcast/cover.png","author":"Fexingo","episode_count":104,"summary":"Lucas and Luna dissect the week in artificial intelligence — not the hype, but the actual models, benchmarks, and deployment decisions shaping the industry. Each episode anchors on a specific paper, product launch, or policy move: from Mixture-of-Experts architecture changes to EU AI Act enforcement, from OpenAI's governance restructuring to open-weight model licensing battles. They compare LLM benchmark scores across reasoning, coding, and multilingual tasks, examine inference cost curves per million tokens, and trace how foundation model competition affects downstream startups. Lucas, a journalist covering tech policy, brings the regulatory and competitive landscape; Luna, an ML engineer turned product lead, presses on technical tradeoffs and real-world performance. Together they avoid speculation and focus on data: what the latest Nvidia GPU cluster means for training efficiency, why a particular transformer variant reduced latency by 40%, or how retrieval-augmented generation changes enterprise search ROI. The show serves engineers, product managers, and investors who need to separate signal from noise in AI. If you want to understand why one billion-parameter model beats anot…","last_synced_at":"2026-07-11T14:18:53.849941+00:00","page_url":"https://stenobird.com/podcast/the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011"},"episode":{"title":"The Hidden Cost of AI Model Inference at Scale","slug":"the-hidden-cost-of-ai-model-inference-at-scale","published_at":"2026-07-02T08:11:05+00:00","page_url":"https://stenobird.com/podcast/the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011/the-hidden-cost-of-ai-model-inference-at-scale","show_page_url":"https://stenobird.com/podcast/the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011","url":"https://audio.fexingo.com/business/the-ai-podcast/episode-0086.mp3","audio_url":"https://audio.fexingo.com/business/the-ai-podcast/episode-0086.mp3","summary":"Lucas and Luna unpack the often-overlooked expense of running AI models in production: inference costs. While training gets the headlines, inference is where the real spending happens for most companies. They explore how inference costs can dwarf training expenses, why hyperscalers are racing to build custom inference chips, and what the recent moves by companies like AMD and Nvidia reveal about the future of AI deployment. The hosts tie in the latest market data, including AMD's 1.6% weekly gain and the broader chip selloff, to show how investors are starting to price in the inference boom. If you're building an AI product or just trying to understand where the money goes in AI, this episode gives you a concrete framework for thinking about the cost of every query. #AIInference #InferenceCosts #AISpending #AMD #NVDA #ChipStocks #Hyperscalers #CustomSilicon #ModelDeployment #MachineLearning #AIChips #TechSpending #FexingoBusiness #BusinessPodcast #TechPodcast #AIEconomy #InferenceHardware #AICostModel Keep every episode free: buymeacoffee.com/fexingo","meta_description":"Lucas and Luna unpack the often-overlooked expense of running AI models in production: inference costs. While training gets the headlines, inference is wh…","key_points":[],"chapters":[],"topics":[],"duration_seconds":528,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011/episodes/the-hidden-cost-of-ai-model-inference-at-scale/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/the-ai-podcast-with-fexingo-artificial-intelligence-machine-learning-and-modern-ai-models-7872011/the-hidden-cost-of-ai-model-inference-at-scale.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}