{"podcast":{"title":"Software Engineering Radio - The Podcast for Professional Software Developers","slug":"software-engineering-radio","podcast_index_feed_id":1372513,"rss_url":"https://feeds.feedburner.com/se-radio","website_url":"https://www.se-radio.net","image_url":"http://media.computer.org/sponsored/podcast/se-radio/se-radio-logo-1400x1475.jpg","author":"SE-Radio Team (team@se-radio.net)","episode_count":724,"summary":"Software Engineering Radio is a podcast targeted at the professional software developer. The goal is to be a lasting educational resource, not a newscast. Every 10 days, a new episode is published that covers all topics software engineering. Episodes are either tutorials on a specific topic, or an interview with a well-known character from the software engineering world. All SE Radio episodes are original content — we do not record conferences or talks given in other venues. Each episode comprises two speakers to ensure a lively listening experience. SE Radio is an independent and non-commercial organization.","last_synced_at":null,"page_url":"https://stenobird.com/podcast/software-engineering-radio"},"episode":{"title":"SE Radio 703: Sahaj Garg on Low Latency AI","slug":"se-radio-703-sahaj-garg-on-low-latency-ai","published_at":"2026-01-14T17:25:00+00:00","page_url":"https://stenobird.com/podcast/software-engineering-radio/se-radio-703-sahaj-garg-on-low-latency-ai","show_page_url":"https://stenobird.com/podcast/software-engineering-radio","url":"https://se-radio.net/2026/01/se-radio-702-derick-schaefer-on-modern-clis/","audio_url":"https://traffic.libsyn.com/secure/seradio/703-sahaj-garg-low-latency-ai.mp3?dest-id=23379","summary":"In this episode, Sahaj Garg, CTO of wispr.ai, joins SE Radio host Robert Blumen to talk about the challenges of building low-latency AI applications. They discuss latency's effect on consumer behavior as well as interactive applications. The conversation explores how to measure latency and how scale impacts it. Then Sahaj and Robert shift to themes around AI, including whether \"AI\" means LLMs or something broader, as they look at latency requirements and challenges around subtypes of AI applications. The final part of the episode explores techniques for managing latency in AI: speed vs accuracy trade-offs; speed vs cost; latency vs cost; choosing the right model; reducing quantization; distillation; and guessing + validating. Brought to you by IEEE Computer Society and IEEE Software magazine.","meta_description":"In this episode, Sahaj Garg, CTO of wispr.ai, joins SE Radio host Robert Blumen to talk about the challenges of building low-latency AI applications. They…","key_points":[],"chapters":[],"topics":[],"duration_seconds":3290,"processing_state":"failed","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/software-engineering-radio/episodes/se-radio-703-sahaj-garg-on-low-latency-ai/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/software-engineering-radio/se-radio-703-sahaj-garg-on-low-latency-ai.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}