الباحثون

Xi Peng

المنشورات 3

نسخة أولية وصول مفتوح

Draft-KV: Learning Useful Latent Communication Between Language Models

Linquan Wu, Shichang Meng, Tianxiang Jiang وآخرون · 2026

Latent communication passes internal states between language models instead of decoded text, but higher receiver accuracy does not show that the receiver used the message content. Across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points, even w …

نسخة أولية وصول مفتوح

Copy the Same, Distill the Difference: Initializing Linear Vision Transformers

Linear Vision Transformers (ViTs) are designed to replace the attention in Softmax ViTs with the linear-complexity attention operator for more efficient token routing, but they require from-scratch pre-training and typically underperform the original Softmax version. How to initialize linear ViTs both efficiently and e …

المؤلفون المشاركون