Abstract
CAD reconstruction methods assume a luxury reality rarely grants: unrestricted visual access to the object, photographed from any desired angle. Real objects, however, are scene-embedded, bolted against walls, wedged into corners, resting on floors, where the scene renders much of the view sphere unreachable and the remaining views unequally informative. We introduce \textbf{SightCAD}, a framework for parametric CAD reconstruction that treats view feasibility as a first-class constraint. In this work we consider objects from standard CAD benchmarks embedded in realistic indoor scenes with physically derived visibility constraints over a discrete view sphere. A learned view selector must choose $K$ feasible views for a vision--language model (VLM) that generates executable CadQuery code, scored by geometric fidelity of the executed solid. Because reward arrives only after discrete view selection, autoregressive generation, and CAD-kernel execution, we propose a joint training paradigm in which the view selector and the CAD-generation VLM are trained together against this reward. The learned selection policy departs sharply from random, uniform, and coverage-greedy alternatives, outperforming surface-area maximization (SA-max) by up to $6.4$ Intersection-over-Union (IoU) points across budgets $K\in\{1,\dots,5\}$. The full system surpasses strong external baselines on scene-embedded, occluded multi-view renders of DeepCAD and Fusion360 objects ($+21$ and $+17$ effective-mIoU points over the best baseline, respectively), as well as on test-time domain-canonicalized real images from the industrial T-LESS benchmark and on both synthetic and real images from the MP6D industrial metal-parts benchmark, while producing the highest rate of executable programs of any method compared (invalid-code rate ${\leq}1.5\%$).
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Bali, K., Guru, M., Borjigin, Y., Starostina, A., Cyron, C. J., & Aydin, R. (2026). Look Where You Can: Active View Selection for CAD Reconstruction under Occlusion. https://omanscience.com/en/articles/look-where-you-can-active-view-selection-for-cad-reconstruction-under-occlusion
MLA 9
Bali, Kartik, et al. "Look Where You Can: Active View Selection for CAD Reconstruction under Occlusion." https://omanscience.com/en/articles/look-where-you-can-active-view-selection-for-cad-reconstruction-under-occlusion.
Chicago (author–date)
Bali, Kartik, Mahish Guru, Yiderigun Borjigin, Alexandra Starostina, Christian J. Cyron, and Roland Aydin. 2026. "Look Where You Can: Active View Selection for CAD Reconstruction under Occlusion." https://omanscience.com/en/articles/look-where-you-can-active-view-selection-for-cad-reconstruction-under-occlusion.
Harvard
Bali, K., Guru, M., Borjigin, Y., Starostina, A., Cyron, C. J. and Aydin, R. (2026) 'Look Where You Can: Active View Selection for CAD Reconstruction under Occlusion', Available at: https://omanscience.com/en/articles/look-where-you-can-active-view-selection-for-cad-reconstruction-under-occlusion.
Vancouver
Bali K, Guru M, Borjigin Y, Starostina A, Cyron CJ, Aydin R. Look Where You Can: Active View Selection for CAD Reconstruction under Occlusion. https://omanscience.com/en/articles/look-where-you-can-active-view-selection-for-cad-reconstruction-under-occlusion
IEEE
K. Bali, M. Guru, Y. Borjigin, A. Starostina, C. J. Cyron, and R. Aydin, "Look Where You Can: Active View Selection for CAD Reconstruction under Occlusion," https://omanscience.com/en/articles/look-where-you-can-active-view-selection-for-cad-reconstruction-under-occlusion.