# Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction Page: https://stenobird.com/podcast/daily-paper-cast-7079649/bridging-videoqa-and-video-guided-agentic-tasks-via-generalized-keyframe-extraction Text version: https://stenobird.com/podcast/daily-paper-cast-7079649/bridging-videoqa-and-video-guided-agentic-tasks-via-generalized-keyframe-extraction.md Podcast: [Daily Paper Cast](https://stenobird.com/podcast/daily-paper-cast-7079649) Published: 2026-07-01T03:27:48+00:00 Episode link: https://share.transistor.fm/s/19acb893 Audio file: https://media.transistor.fm/19acb893/0b26da1d.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/bridging-videoqa-and-video-guided-agentic-tasks-via-generalized-keyframe-extraction Duration seconds: 1370 ## Resource 🤗 Upvotes: 22 | cs.CV, cs.AI Authors: Sunqi Fan, Qingle Liu, Runqi Yin, Meng-Hao Guo, Shuojin Yang Title: Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction Arxiv: http://arxiv.org/abs/2606.29445v1 Abstract: Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achieved remarkable performance on Video Question Answering (VideoQA) benchmarks. However, existing benchmarks primarily evaluate whether models can perceive shallow visual cues, while rarely examining whether MLLMs can learn deeper knowledge or procedural skills from video tutorials and generalize them to downstream long-horizon agentic tasks. To address this gap, we introduce VG-GUIBench (Video-Guided GUI Benchmark), a new benchmark designed to evaluate whether MLLM-based GUI agents can follow video tutorials to complete corresponding GUI interactive tasks. Furthermore, we observe that the performance of models on both VideoQA and video-guided agentic tasks critically depends on effective keyframe extraction. Based on this observation, we propose TASKER (Task-driven And Scene-aware Keyframe searchER), a keyframe extraction algorithm that jointly considers task relevance and scene dynamics to identify informative frames. Experimental results demonstrate that TASKER achieves significant performance improvements on both VideoQA and video-guided agentic task benchmarks, outperforming the best baseline by 2.0% on the EgoSchema fullset and 1.8% on the NExT-QA dataset, respectively. These results further highlight the potential of generalized keyframe extraction methods for video understanding tasks. Our code and data are available at https://github.com/VG-GUI-TASKER/VG-GUI-TASKER. ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/bridging-videoqa-and-video-guided-agentic-tasks-via-generalized-keyframe-extraction/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/daily-paper-cast-7079649/bridging-videoqa-and-video-guided-agentic-tasks-via-generalized-keyframe-extraction.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.