الباحثون

Federico Tombari

المنشورات 4

نسخة أولية وصول مفتوح

Level-of-Token Diffusion

Kiyohiro Nakayama, Brian Chao, Jan Ackermann وآخرون · 2026

Image and video diffusion models allocate equal computation to every region, even when the intended scene calls for varying levels of detail. The spatial distribution of detail can often be anticipated before generation, indicating where computation can be reduced. We introduce Level-of-Token (LoT) Diffusion, a framewo …

نسخة أولية وصول مفتوح

ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning

Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This requires connecting 4D interaction understanding, which explains how past actions changed the scene, with spatially grounded planning, which determines how and where to …

نسخة أولية وصول مفتوح

LEGAU: Learning Semantic Gaussian Priors for Scalable Category-level Pose Estimation

Hongli Xu, Zhaowei Lu, Junwen Huang وآخرون · 2026

Category-level 6D pose estimation from a single RGB-D observation is inherently under-constrained, since partial visible geometry must be interpreted together with a canonical object structure before a stable pose can be determined. We present LEGAU, a unified framework that jointly predicts NOCS correspondence, object …

المؤلفون المشاركون