الباحثون

Xin Lu

المنشورات 3

نسخة أولية وصول مفتوح

ScribbleEdit: A Benchmark for Scribble-Only Image Editing

Jie Ren, Hao Kang, Kai Guo وآخرون · 2026

Scribble-based interaction provides a lightweight and intuitive way for users to specify image editing intents in interactive editing tools. However, current image editing models based on VLMs or LLMs struggle to understand and execute edits based solely on scribble inputs. To systematically study this problem, we cons …

نسخة أولية وصول مفتوح

DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models

Haojun Xu, Jie Huang, Xin Lu وآخرون · 2026

Large video diffusion models offer expressive priors for embodied prediction and learning, yet their many-step sampling remains costly for interactive downstream use. Distribution Matching Distillation (DMD) enables few-step video generation, but can suppress robot--object motion while preserving visual quality. Examin …

نسخة أولية وصول مفتوح

Beyond Spatial Benchmarks: From Spatial Reasoning to Navigation

Xun Huang, Shijia Zhao, Rongsheng Qu وآخرون · 2026

Does progress on spatial reasoning benchmarks translate into better navigation? Existing benchmarks test isolated inferences from images or videos, with little connection to downstream navigation. Our analysis reveals a gap between benchmark-oriented spatial specialization and navigation performance, and shows how alig …

المؤلفون المشاركون