الباحثون

Chi-Wing Fu

المنشورات 3

نسخة أولية وصول مفتوح

Seek-and-View Reasoning for Multi-View Spatial Understanding

Qixiang Chen, Cheng Zhang, Fucai Ke وآخرون · 2026

Existing approaches to multi-view spatial reasoning operate largely on sparse input views. Vision-language models (VLMs) are thus restricted to understand a scene and infer spatial relations within these fixed views, leading to fragile cross-view alignment and geometry-to-language bottleneck. To address these issues, w …

نسخة أولية وصول مفتوح

ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing

Donghao Zhou, Haoyang He, Fan Zhang وآخرون · 2026

Instruction-guided video editing has made significant progress, yet existing methods use multimodal large language models (MLLMs) primarily as semantic encoders, so they often fall short in working with implicit edits that require causal or semantic reasoning. To bridge this fundamental gap in video editing, we propose …

نسخة أولية وصول مفتوح

PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing

Donghao Zhou, Jia-Hui Pan, Fan Zhang وآخرون · 2026

Robotic bin packing requires long-horizon sequential decision-making, as each object placement affects the available space for subsequent packing. Existing methods primarily rely on hand-crafted geometric heuristics that optimize predefined objectives or reinforcement learning policies learned through trial and error o …

المؤلفون المشاركون