Abstract
Active robot vision requires controlling the camera to reveal task-relevant information that is hidden from the current viewpoint. For example, determining what is inside a box may require raising the camera and looking down into it. For viewpoint-dependent question answering, the challenge is to select camera motions that expose the visual evidence needed to answer the question. Although vision-language models (VLMs) can interpret observed images, selecting such motions requires anticipating the usefulness of unseen views. We quantify this usefulness as answerability, a VLM's estimate that a view suffices to answer the question, and present Rendering-Free Lookahead (RFL), a viewpoint-selection policy that ranks candidate camera motions by predicted future answerability. RFL transfers visual lookahead from deployment to offline training. At training, a privileged teacher renders candidate future views in 3D Gaussian Splatting (3DGS) scenes and uses a frozen VLM to compute one- and two-step answerability targets. Through two-stage distillation, a student learns to predict these action values from the question, recent visual observations, and a candidate camera motion. At deployment, RFL uses these predicted values to select camera motions without rendering future views. On 377 E3VS-Bench test episodes in unseen environments, RFL improves the mean judge score by 43\% over a direct-action baseline using the same VLM. These results support learning camera-control policies from privileged visual lookahead for viewpoint-dependent question answering.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Sakamoto, K., Azuma, D., Kurita, S., Chiba, N., Iwasawa, Y., Matsuo, Y., & Miyanishi, T. (2026). Rendering-Free Lookahead for Question-Guided Active Vision. https://omanscience.com/en/articles/rendering-free-lookahead-for-question-guided-active-vision
MLA 9
Sakamoto, Koya, et al. "Rendering-Free Lookahead for Question-Guided Active Vision." https://omanscience.com/en/articles/rendering-free-lookahead-for-question-guided-active-vision.
Chicago (author–date)
Sakamoto, Koya, Daichi Azuma, Shuhei Kurita, Naoya Chiba, Yusuke Iwasawa, Yutaka Matsuo, and Taiki Miyanishi. 2026. "Rendering-Free Lookahead for Question-Guided Active Vision." https://omanscience.com/en/articles/rendering-free-lookahead-for-question-guided-active-vision.
Harvard
Sakamoto, K., Azuma, D., Kurita, S., Chiba, N., Iwasawa, Y., Matsuo, Y. and Miyanishi, T. (2026) 'Rendering-Free Lookahead for Question-Guided Active Vision', Available at: https://omanscience.com/en/articles/rendering-free-lookahead-for-question-guided-active-vision.
Vancouver
Sakamoto K, Azuma D, Kurita S, Chiba N, Iwasawa Y, Matsuo Y, et al. Rendering-Free Lookahead for Question-Guided Active Vision. https://omanscience.com/en/articles/rendering-free-lookahead-for-question-guided-active-vision
IEEE
K. Sakamoto, D. Azuma, S. Kurita, N. Chiba, Y. Iwasawa, Y. Matsuo, and T. Miyanishi, "Rendering-Free Lookahead for Question-Guided Active Vision," https://omanscience.com/en/articles/rendering-free-lookahead-for-question-guided-active-vision.