# Beacon: Knowing When and How to Perform Agentic Visual Reasoning Page: https://stenobird.com/podcast/daily-paper-cast-7079649/beacon-knowing-when-and-how-to-perform-agentic-visual-reasoning Text version: https://stenobird.com/podcast/daily-paper-cast-7079649/beacon-knowing-when-and-how-to-perform-agentic-visual-reasoning.md Podcast: [Daily Paper Cast](https://stenobird.com/podcast/daily-paper-cast-7079649) Published: 2026-08-01T03:53:05+00:00 Episode link: https://share.transistor.fm/s/11578e60 Audio file: https://media.transistor.fm/11578e60/cfdb1769.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/beacon-knowing-when-and-how-to-perform-agentic-visual-reasoning Duration seconds: 1218 ## Resource 🤗 Upvotes: 45 | cs.CV Authors: Qixun Wang, Yang Shi, Letian Cheng, Zhuoran Zhang, Yan He, Yuqi Tang, Qi Zhang, Xinlei Yu, Ruizhe Chen, Tianrun Xu, Yuanxing Zhang, Pengfei Wan, Haotian Wang, Xianghua Ying Title: Beacon: Knowing When and How to Perform Agentic Visual Reasoning Arxiv: http://arxiv.org/abs/2607.28595v1 Abstract: The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet inefficient reasoning paradigm. In this work, we rethink agentic visual reasoning through two key dimensions of tool use: Mode Adaptiveness (MA) and Tool Effect (TE). Mode Adaptiveness characterizes whether an MLLM can recognize when tools are truly necessary and invoke them accordingly, thereby avoiding unnecessary computational overhead while improving performance on challenging problems that require tool assistance. Tool Effect characterizes the actual impact of tool use: tools should extend the model's capabilities on problems unsolvable through text-only reasoning, while avoiding additional errors on problems that the model can already solve without tools. We conduct a comprehensive analysis to quantify these two properties and empirically reveal that existing agentic visual reasoning models exhibit limited Mode Adaptiveness, while the gains produced by tool use on hard examples are largely offset by the harm introduced on easy examples that the models can already solve. Motivated by these observations, we propose Beacon, a novel agentic visual reasoning model that achieves stronger overall performance, improved Mode Adaptiveness, and genuine tool-induced performance gains. At the core of Beacon are the Necessity-Aware Adaptive Reward and the Hin… ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/beacon-knowing-when-and-how-to-perform-agentic-visual-reasoning/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/daily-paper-cast-7079649/beacon-knowing-when-and-how-to-perform-agentic-visual-reasoning.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.