نسخة أولية وصول مفتوح
Quasi Linear Kernel Attention with Infinite Capacity
The evaluation cost of transformers with softmax attention scales quadratically with sequence length. Kernel attention addresses this by replacing softmax with a more general kernel function. In this paper, we aim to identify kernels that retain the expressivity of attention while enabling quasi linear computation. To …