نسخة أولية وصول مفتوح
Scaling Parameter and Context in Attention: Native Sparse Attention from Mixture-of-Head
Scaling attention parameters can improve language model quality, but retaining full token histories makes additional heads costly at long contexts. Furthermore, since attention retrieves and combines contextual information, parameter scaling should also support longer contexts. We therefore ask whether attention parame …