الباحثون

Jiayin Zheng

المنشورات 1

نسخة أولية وصول مفتوح

IronViT: Toward Efficient Generalist Visual Representation Learning

Jiaxi Huang, Yueqi Hu, Xin Zhu وآخرون · 2026

A generalist vision encoder must capture semantic, spatial, language-aligned, and action-relevant cues within a unified representation, yet softmax attention underlying today's most capable visual backbones becomes prohibitively expensive at high resolution. A natural attempt to address both challenges is to distill mu …

المؤلفون المشاركون