الباحثون

Alan Yuille

المنشورات 5

نسخة أولية وصول مفتوح

Transforming Image Editors into Video Editors

Feng Wang, Zijie Li, Ceyuan Yang وآخرون · 2026

Recent image editing systems have achieved impressive semantic understanding, visual fidelity, and instruction-following ability, while video editing remains substantially more difficult and costly. In this paper, we present a simple alternative to end-to-end video editing: instead of training a monolithic video editor …

نسخة أولية وصول مفتوح

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

Ruihong Shen, Žiga Kovačič, Peter Kulits وآخرون · 2026

We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visual observations into compact representations of scene structure and dynamics, by implementing abstra …

نسخة أولية وصول مفتوح

Generative Cinematographer: Composing Camera and Object Motion in 3D

Current controllable video generation systems often rely on 2D motion trajectories or sparse drag signals for object motion. These controls are ambiguous because the same 2D trajectory can correspond to different 3D motions, especially when the camera and objects move simultaneously. We present Generative Cinematograph …

نسخة أولية وصول مفتوح

FurE: Efficient Instance-Specific 3D Fur Reconstruction without Animal-Fur Datasets

Realistic and editable animal fur reconstruction from multi-view images is challenging due to fine-scale detail, self-occlusion and obfuscation, and, unlike human hair, the lack of animal-fur datasets. Fur usually covers most of an animal's body, with large inter-species and intra-species variability. We present FurE, …

نسخة أولية وصول مفتوح

Training Object Permanence in World Models

Haotian Zhang, Fengyuan Yu, Dezhi Luo وآخرون · 2026

Object permanence and solidity are hallmarks of human cognitive priors. Recent studies show that video generation models, a paradigmatic class of current world models, have begun to show emerged reasoning abilities, making them ideal candidates for building human-like physical intelligence. Do video models have emerged …

المؤلفون المشاركون