الباحثون

Huaiyuan Qin

المنشورات 1

نسخة أولية وصول مفتوح

Copy the Same, Distill the Difference: Initializing Linear Vision Transformers

Linear Vision Transformers (ViTs) are designed to replace the attention in Softmax ViTs with the linear-complexity attention operator for more efficient token routing, but they require from-scratch pre-training and typically underperform the original Softmax version. How to initialize linear ViTs both efficiently and e …

المؤلفون المشاركون