الباحثون

Xiaochuang Han

المنشورات 3

نسخة أولية وصول مفتوح

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie وآخرون · 2026

Training a model to generate visual content can encourage it to learn rich perceptual capabilities related to geometry, spatial relationships, and objectness; yet, its benefits for visual understanding remain unclear. We ask: when and how does visual generation supervision improve visual understanding? We study control …

نسخة أولية وصول مفتوح

Verifiable Visual Rewards Transfer from Synthetic Scenes to Natural Prompts

Precise instruction following in image generation, such as satisfying object counts and spatial relations, remains an open challenge at least in part because it is learned using unreliable reward models such as object detectors and vision-language models. We introduce Verifiable Visual Rewards (VVR), the first framewor …

المؤلفون المشاركون