الباحثون

Zhibo Chen

المنشورات 5

نسخة أولية وصول مفتوح

VAMR: Multi-Question Agentic Reasoning for Efficient Long-Form Video Understanding

Runquan Gui, Hanzhu Chen, Zehao Wang وآخرون · 2026

Long-form video understanding often involves multiple questions about different aspects of the same recording. Yet existing video agents typically process each question through an isolated tool-use trajectory. This repeatedly restarts video exploration and memory construction, missing opportunities to acquire evidence …

نسخة أولية وصول مفتوح

PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop

Xinge Peng, Yiting Lu, Tianwu Zhi وآخرون · 2026

Vision-Language Models (VLMs) have shown strong multimodal reasoning capabilities, yet whether they truly capture the physical consistency underlying real-world dynamics remains unclear. Existing benchmark paradigms often suffer from fragmented evaluation, focusing on isolated cognitive stages while overlooking the inh …

نسخة أولية وصول مفتوح

ReCaVSR: One-Step Streaming Diffusion Video Super-Resolution with Recycled Latents and Learned Cache Routing

Xijun Wang, Xin Li, Suhang Yao وآخرون · 2026

Real-time diffusion-based video super-resolution (VSR) is in high demand for online streaming, yet stringent latency requirements often compromise generative fidelity. We propose ReCaVSR, a Wan2.2-based, one-step framework for streaming VSR that builds on two observations: recycled SR latents retain local temporal cont …

نسخة أولية وصول مفتوح

RelayVSR: Large-Small Model Collaboration for Efficient Real-World Video Super-Resolution

Xijun Wang, Xin Li, Zirui Lang وآخرون · 2026

Large generative models can recover realistic detail in real-world video super-resolution (VSR), but processing an entire video with them is computationally expensive. In this work, we present RelayVSR, a streaming VSR framework built on the Sparse Generative Relay mechanism. A large generative model generates referenc …

نسخة أولية وصول مفتوح

CoDrive: Cross-Vehicle World-Consistent Video Generation with Precise Trajectory Control for Cooperative Driving

Yu Meng, Baining Zhao, Junta Wu وآخرون · 2026

Real-world driving is inherently multi-agent, yet most existing driving world models generate observations from a single ego vehicle. Independently extending them to multiple vehicles does not ensure that different agents observe a consistent shared world. We present CoDrive, a cross-vehicle, multi-view driving video g …

المؤلفون المشاركون