الباحثون

Yadan Luo

المنشورات 4

نسخة أولية وصول مفتوح

VLM4Cluster: Benchmarking Deep Clustering In the Era of Vision-Language Pre-training

Yuanwei Hu, Bo Peng, Yuheng Jia وآخرون · 2026

Vision-language pre-training has reshaped image clustering, giving rise to language-assisted image clustering (LaIC), which leverages textual semantics to complement visual representations. Despite the rapid proliferation of LaIC methods, it remains unclear how much LaIC has actually advanced image clustering, as exist …

نسخة أولية وصول مفتوح

Does the VGGT Family Need All Its Layers?

Fengyi Zhang, Holger Caesar, Xiangyu Sun وآخرون · 2026

Which layers of a feed-forward geometry model are needed to preserve both camera poses and dense 3D structure? We study layer redundancy in VGGT, $π^3$, and VGGT-$Ω$: 3,018 pruned configurations, scored on seven camera-pose and dense-geometry metrics across four indoor and outdoor datasets. Four findings follow: (i) Re …

نسخة أولية وصول مفتوح

Train Together or Merge Later? Unifying VLA Experts via a Shared Action Interface

Zhizhen Zhang, Yuxia Fu, Zijian Wang وآخرون · 2026

Co-training offers a straightforward way to build a multi-task vision-language-action (VLA) policy, but can fall short of the performance achieved by training each task independently. The challenge is to retain these task-specific gains in a multi-task policy without joint post-training. Combining independently trained …

نسخة أولية وصول مفتوح

Open Vocabulary Domain Unlearning

Sumanth Udupa, Mehrtash Harandi, Yadan Luo وآخرون · 2026

Vision-Language Models (VLMs) exhibit remarkable zero-shot generalization, yet they often encode unwanted or hazardous stylistic domains such as idealized textbook diagrams in medical AI or cartoon vehicles in autonomous driving. Approximate Domain Unlearning (ADU) aims to selectively erase a model's recognition of a t …

المؤلفون المشاركون