{"podcast":{"title":"Vanishing Gradients","slug":"vanishing-gradients-4989163","podcast_index_feed_id":4989163,"rss_url":"https://api.substack.com/feed/podcast/2632531.rss","website_url":"https://hugobowne.substack.com/podcast","image_url":"https://substackcdn.com/feed/podcast/2632531/e8d57d9d781f20857949c2678ef8c9c2.jpg","author":"Hugo Bowne-Anderson","episode_count":77,"summary":"a data podcast with hugo bowne-anderson","last_synced_at":null,"page_url":"https://stenobird.com/podcast/vanishing-gradients-4989163"},"episode":{"title":"LLM Architecture in 2026: What You Need to Know with Sebastian Raschka","slug":"llm-architecture-in-2026-what-you-need-to-know-with-sebastian-raschka","published_at":"2026-04-13T02:52:27+00:00","page_url":"https://stenobird.com/podcast/vanishing-gradients-4989163/llm-architecture-in-2026-what-you-need-to-know-with-sebastian-raschka","show_page_url":"https://stenobird.com/podcast/vanishing-gradients-4989163","url":"https://hugobowne.substack.com/p/llm-architecture-in-2026-what-you","audio_url":"https://api.substack.com/feed/podcast/194026308/c9e012a9f1d548c39da8bd29e7514798.mp3","summary":"If you take a model release as an anchor point , let’s say Nemotron 3 or Qwen 3.5, you can go in both directions : You can either plug them into an agent and play around with that, or you can look, okay, what does the model look like under the hood ? What are the ingredients? What type of attention mechanism do they use? What are currently research techniques that could make that even better in the next generation of models? What can we swap out, basically? And I’m interested in both of these! Sebastian Raschka , Independent AI Researcher and author of Build a Large Language Model from Scratch , joins Hugo to talk about what’s changed in AI architecture, from post-training to hybrid models, and why understanding what’s under the hood matters more than ever for developers building in the agentic era . Sebastian’s upcoming book, Build a Reasoning Model from Scratch , currently available for pre-order on Amazon and in early access on Manning ! Vanishing Gradients is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. We Discuss: * Ed Tech for Agents : should we design educational content specifically for agentic systems, or is there a better approach? * Inference Scaling is the new frontier, driving “gold-level” performance during generation via parallel sampling and internal meta-judges ; * Hybrid Architectures from Qwen 3.5 and Nemotron 3 scale almost linearly, making long-context agentic workflows significantly more affordable and performant; * Multi-head Latent Attention (MLA) , developed by DeepSeek , wins the KV cache war by drastically reducing memory overhead without performance hits; * Agent Harnesses need to be continuously simplified as frontier models are post-trained on agent trajectories. Tea…","meta_description":"If you take a model release as an anchor point , let’s say Nemotron 3 or Qwen 3.5, you can go in both directions : You can either plug them into an agent…","key_points":[],"chapters":[],"topics":[],"duration_seconds":4682,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/vanishing-gradients-4989163/episodes/llm-architecture-in-2026-what-you-need-to-know-with-sebastian-raschka/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/vanishing-gradients-4989163/llm-architecture-in-2026-what-you-need-to-know-with-sebastian-raschka.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}