الباحثون

Chang Xu

المنشورات 9

نسخة أولية وصول مفتوح

VibeEdit: Image Editing with Canvas Instructions

Jinjing Zhao, Fangyun Wei, Yitong Wang وآخرون · 2026

In text-guided image editing, describing the desired change is often straightforward, but identifying the intended object or region can be cumbersome, especially when several objects look alike. We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image …

نسخة أولية وصول مفتوح

VDOT++: Unified Few-Step Video Generation via Unbalanced Optimal Transport Distillation

Yutong Wang, Xingtong Ge, Enhuai Liu وآخرون · 2026

Video creation spans text-to-video (T2V), image-to-video (I2V), and condition-based generation, yet video diffusion models remain costly because they repeatedly evaluate large backbones during sampling. Distribution matching distillation (DMD) reduces this cost, but its reverse Kullback--Leibler (KL) objective can prov …

نسخة أولية وصول مفتوح

Revealing Epistemic Uncertainty in MLLMs via Causal-Invariant Masking

Haoyang Luo, Linwei Tao, Jie Gui وآخرون · 2026

Multimodal Large Language Models (MLLMs) suffer from hallucinations, creating a critical need for Uncertainty Quantification (UQ) to ensure reliable deployment. However, existing approaches struggle to detect uncertainty caused by superficial associations, especially when the query-relevant signal is weak. We mainly at …

نسخة أولية وصول مفتوح

ShelfChange3D: Object-Level 3D Change Detection for Retail Shelf Monitoring

Lingyi Zhou, Yunke Wang, Mengyu Zheng وآخرون · 2026

Reliable shelf monitoring is an important capability for retail automation, yet existing out-of-stock detection methods mainly operate in image space and lack metric 3D localization for downstream robotic systems. We formulate shelf monitoring as object-level 3D change detection: given two RGB-D observations captured a …

نسخة أولية وصول مفتوح

DSDyn-VLA: A Dual-Stream Dynamic Manipulation Framework with Motion Perception, Future Awareness, and Realtime Correction

Wenhao Li, Xiu Su, Yu Han وآخرون · 2026

While Vision-Language-Action (VLA) models excel in static tasks, they struggle in dynamic environments where objects are in motion (e.g., conveyor belt manipulation). We identify three fundamental limitations hindering current VLAs in these scenarios: the \textbf{perception gap}, where static visual inputs lack tempora …

نسخة أولية وصول مفتوح

PulseQuant: Propagation-Guided Subspace Correction for 4-Bit Video Diffusion Transformers

Yutong Wang, Xingtong Ge, Enhuai Liu وآخرون · 2026

Quantization errors in video diffusion transformers can be amplified or attenuated by subsequent denoising updates, making local reconstruction error an incomplete predictor of final impact. We introduce PulseQuant, a 4-bit post-training quantization method that combines trajectory sensitivity with activation geometry …

المؤلفون المشاركون