# Building AI Voice Agents with Scott Stephenson - #707 Page: https://stenobird.com/podcast/twiml-ai-podcast/building-ai-voice-agents-with-scott-stephenson-707 Text version: https://stenobird.com/podcast/twiml-ai-podcast/building-ai-voice-agents-with-scott-stephenson-707.md Podcast: [The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)](https://stenobird.com/podcast/twiml-ai-podcast) Published: 2024-10-28T16:36:00+00:00 Episode link: https://twimlai.com/podcast/twimlai/building-ai-voice-agents/ Audio file: https://pscrb.fm/rss/p/traffic.megaphone.fm/MLN6815731992.mp3?updated=1730147923 Processing state: failed JSON: https://stenobird.com/v1/public/podcasts/twiml-ai-podcast/episodes/building-ai-voice-agents-with-scott-stephenson-707 Duration seconds: 3704 ## Resource Today, we're joined by Scott Stephenson, co-founder and CEO of Deepgram to discuss voice AI agents. We explore the importance of perception, understanding, and interaction and how these key components work together in building intelligent AI voice agents. We discuss the role of multimodal LLMs as well as speech-to-text and text-to-speech models in building AI voice agents, and dig into the benefits and limitations of text-based approaches to voice interactions. We dig into what’s required to deliver real-time voice interactions and the promise of closed-loop, continuously improving, federated learning agents. Finally, Scott shares practical applications of AI voice agents at Deepgram and provides an overview of their newly released agent toolkit. The complete show notes for this episode can be found at https://twimlai.com/go/707. ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/twiml-ai-podcast/episodes/building-ai-voice-agents-with-scott-stephenson-707/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/twiml-ai-podcast/building-ai-voice-agents-with-scott-stephenson-707.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.