الباحثون

Zi Huang

المنشورات 2

نسخة أولية وصول مفتوح

Does the VGGT Family Need All Its Layers?

Fengyi Zhang, Holger Caesar, Xiangyu Sun وآخرون · 2026

Which layers of a feed-forward geometry model are needed to preserve both camera poses and dense 3D structure? We study layer redundancy in VGGT, $π^3$, and VGGT-$Ω$: 3,018 pruned configurations, scored on seven camera-pose and dense-geometry metrics across four indoor and outdoor datasets. Four findings follow: (i) Re …

نسخة أولية وصول مفتوح

Tracing the Evidence: Faithful Token Attribution Through Vision-Language Reasoning

Bowen Yuan, Danny Wang, Ruihong Qiu وآخرون · 2026

Large vision-language models (LVLMs) exhibit strong reasoning capabilities, yet the visual and textual evidence supporting the generated responses remains difficult to identify. Faithful token attribution explains an LVLM's response by assigning scores that rank image and prompt tokens by how much the model relies on t …

المؤلفون المشاركون