الباحثون

Gang Xu

المنشورات 4

نسخة أولية وصول مفتوح

Juno: Taming Predictive Latents for Vision-Language-Action Models

Yuchen Zhu, Chenyi Xu, Yulin Zhang وآخرون · 2026

Joint-embedding predictive architectures (JEPAs) predict masked or future observations in representation space, offering a natural source of predictive latents for vision-language-action (VLA) models. Yet making these latents useful across pretraining, policy learning, and deployment requires addressing three failures: …

نسخة أولية وصول مفتوح

Casual Flash Lighting for Gaussian Splat Inverse Rendering

Jiamin Xu, Dongheng Wei, Jiarong Zhao وآخرون · 2026

Recovering geometry, materials, and lighting from photographs is highly ambiguous when only static illumination is available. Active-lighting setups reduce the ambiguity but require dark rooms or specialized hardware. Instead, we synergize both static and flash lighting from casual indoor capture, with the flash on or …

نسخة أولية وصول مفتوح

Can Vision-Language Models Stay Helpful When Facing Implicit Risks? Intent-Privilege OPSD for Efficient Safety-Helpfulness Alignment

Haotian Deng, Wenbin Xing, Gang Xu وآخرون · 2026

Vision-Language Models (VLMs) remain vulnerable to cross-modal implicit risks: visual and textual inputs that appear benign in isolation can jointly elicit unsafe responses. Existing safety methods often require large preference datasets, costly multi-rollout training, or additional safeguards at inference time. They m …

المؤلفون المشاركون