{"podcast":{"title":"Unsupervised Learning with Jacob Effron","slug":"unsupervised-learning-with-jacob-effron-6041643","podcast_index_feed_id":6041643,"rss_url":"https://feeds.simplecast.com/dOSE_bdP","website_url":"https://unsupervised-learning.simplecast.com","image_url":"https://image.simplecastcdn.com/images/ff0dbf2f-5711-4964-9172-807c39ca4824/be7895b2-4fcc-4f70-9864-584710596b1d/3000x3000/redpoint-unsupervised-learning-logo.jpg?aid=rss_feed","author":"Redpoint Ventures","episode_count":98,"summary":"We probe the sharpest minds in AI in search for the truth about what’s real today, what will be real in the future and what it all means for businesses and the world. If you’re a builder, researcher or investor navigating the AI world, this podcast will help you deconstruct and understand the most important breakthroughs and see a clearer picture of reality. Follow this show and consider enabling notifications to stay up to date on our latest episodes. Unsupervised Learning is a podcast by Redpoint Ventures, an early-stage venture capital fund that has invested in companies like Snowflake, Stripe, and Mistral. Hosted by Redpoint investor Jacob Effron alongside Patrick Chase, Jordan Segall and Erica Brescia.","last_synced_at":"2026-07-10T00:20:35.953212+00:00","page_url":"https://stenobird.com/podcast/unsupervised-learning-with-jacob-effron-6041643"},"episode":{"title":"Ep 74: Chief Scientist of Together.AI Tri Dao On The End of Nvidia's Dominance, Why Inference Costs Fell & The Next 10X in Speed","slug":"ep-74-chief-scientist-of-together-ai-tri-dao-on-the-end-of-nvidia-s-dominance-why-inference-costs-fell-the-next-10x-in-speed","published_at":"2025-09-10T12:50:19+00:00","page_url":"https://stenobird.com/podcast/unsupervised-learning-with-jacob-effron-6041643/ep-74-chief-scientist-of-together-ai-tri-dao-on-the-end-of-nvidia-s-dominance-why-inference-costs-fell-the-next-10x-in-speed","show_page_url":"https://stenobird.com/podcast/unsupervised-learning-with-jacob-effron-6041643","url":"https://unsupervised-learning.simplecast.com/episodes/ep-74-chief-scientist-of-togetherai-tri-dao-on-ai-super-researcher-the-end-of-nvidias-dominance-why-inference-costs-fell-the-next-10x-in-speed-hqhgrAYK","audio_url":"https://cdn.simplecast.com/audio/2c08ad29-5b79-42c0-a40a-6c1af4327f2f/episodes/4e1ea3f1-9c21-4670-9ee8-75a1d636a67a/audio/8ea58360-ea33-4b72-9c64-14a609d832b2/default_tc.mp3?aid=rss_feed&feed=dOSE_bdP","summary":"Tri Dao, Chief Scientist at Together AI and Princeton professor who created Flash Attention and Mamba, discusses how inference optimization has driven costs down 100x since ChatGPT's launch through memory optimization, sparsity advances, and hardware-software co-design. He predicts the AI hardware landscape will shift from Nvidia's current 90% dominance to a more diversified ecosystem within 2-3 years, as specialized chips emerge for distinct workload categories: low-latency agentic systems, high-throughput batch processing, and interactive chatbots. Dao shares his surprise at AI models becoming genuinely useful for expert-level work, making him 1.5x more productive at GPU kernel optimization through tools like Claude Code and O1. The conversation explores whether current transformer architectures can reach expert-level AI performance or if approaches like mixture of experts and state space models are necessary to achieve AGI at reasonable costs. Looking ahead, Dao sees another 10x cost reduction coming from continued hardware specialization, improved kernels, and architectural advances like ultra-sparse models, while emphasizing that the biggest challenge remains generating expert-level training data for domains lacking extensive internet coverage.","meta_description":"Tri Dao, Chief Scientist at Together AI and Princeton professor who created Flash Attention and Mamba, discusses how inference optimization has driven cos…","key_points":[],"chapters":[],"topics":[],"duration_seconds":3517,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/unsupervised-learning-with-jacob-effron-6041643/episodes/ep-74-chief-scientist-of-together-ai-tri-dao-on-the-end-of-nvidia-s-dominance-why-inference-costs-fell-the-next-10x-in-speed/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/unsupervised-learning-with-jacob-effron-6041643/ep-74-chief-scientist-of-together-ai-tri-dao-on-the-end-of-nvidia-s-dominance-why-inference-costs-fell-the-next-10x-in-speed.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}