{"podcast":{"title":"The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)","slug":"twiml-ai-podcast","podcast_index_feed_id":1045879,"rss_url":"https://feeds.megaphone.fm/MLN2155636147","website_url":"https://twimlai.com","image_url":"https://megaphone.imgix.net/podcasts/35230150-ee98-11eb-ad1a-b38cbabcd053/image/TWIML_AI_Podcast_Official_Cover_Art_1400px.png?ixlib=rails-4.3.1&max-w=3000&max-h=3000&fit=crop&auto=format,compress","author":"TWIML","episode_count":785,"summary":"Machine learning and artificial intelligence are dramatically changing the way businesses operate and people live. The TWIML AI Podcast brings the top minds and ideas from the world of ML and AI to a broad and influential community of ML/AI researchers, data scientists, engineers and tech-savvy business and IT leaders. Hosted by Sam Charrington, a sought after industry analyst, speaker, commentator and thought leader. Technologies covered include machine learning, artificial intelligence, deep learning, natural language processing, neural networks, analytics, computer science, data science and more.","last_synced_at":null,"page_url":"https://stenobird.com/podcast/twiml-ai-podcast"},"episode":{"title":"Proactive Agents for the Web with Devi Parikh - #756","slug":"proactive-agents-for-the-web-with-devi-parikh-756","published_at":"2025-11-19T01:49:00+00:00","page_url":"https://stenobird.com/podcast/twiml-ai-podcast/proactive-agents-for-the-web-with-devi-parikh-756","show_page_url":"https://stenobird.com/podcast/twiml-ai-podcast","url":"https://twimlai.com/podcast/twimlai/proactive-agents-for-the-web/","audio_url":"https://pscrb.fm/rss/p/traffic.megaphone.fm/MLN8999995371.mp3?updated=1763502496","summary":"The future of web interaction lies in moving from manual clicking to high-level abstraction via proactive, autonomous agents. Devi Parikh explains how Yutori uses visually-grounded models to navigate the web more reliably than traditional DOM-based approaches.","meta_description":"Explore the shift from manual web browsing to proactive AI agents with Yutori co-CEO Devi Parikh. Learn about vision-based web navigation and automation.","key_points":["Main idea: Moving from DOM-based parsing to vision-based models provides much higher robustness against brittle web interfaces","Technical approach: Yutori utilizes a training pipeline involving supervised fine-tuning, rejection sampling, and reinforcement learning","Practical takeaway: Using 'Scouts' allows for ambient, background automation that monitors the web and reports findings without active user input","Failure mode: Traditional browser automation often breaks due to edge cases in website structures, necessitating a shift toward visual grounding","Future vision: The goal is to transition from simple information monitoring to complex, multi-step task automation that operates autonomously"],"chapters":[{"start_ms":60000,"title":"The Evolution of Web Interaction","summary":"A look back at the progress in AI and the shift toward browser-use agents."},{"start_ms":555000,"title":"The Rise of Browser Agents","summary":"Discussing the excitement around automating web tasks and the potential for broader platforms."},{"start_ms":1325000,"title":"Scaling Complex Workflows","summary":"How improving foundation models and custom training pipelines pushes the ceiling of agent capabilities."},{"start_ms":1780000,"title":"Beyond Static Reports","summary":"Moving from simple data retrieval to interactive, actionable outputs from web agents."},{"start_ms":2260000,"title":"The Shift to Vision-Based Navigation","summary":"Why relying on screenshots and visual grounding is more reliable than parsing the DOM."},{"start_ms":2785000,"title":"Adaptive Orchestration","summary":"How 'Scouts' use adaptive plans and tool-use to execute complex, multi-step web tasks."},{"start_ms":3030000,"title":"Ambient Agentic Systems","summary":"The concept of background agents that monitor the web 24/7 and notify users of significant events."}],"topics":["Proactive Agents","Web Automation","Computer Vision","Multimodal Models","Browser Use Models","Autonomous Agents","Yutori","AI Agents"],"duration_seconds":3364,"processing_state":"processed","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/twiml-ai-podcast/episodes/proactive-agents-for-the-web-with-devi-parikh-756/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/twiml-ai-podcast/proactive-agents-for-the-web-with-devi-parikh-756.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}