الباحثون

Thomas Vaitses Fontanari

المنشورات 1

نسخة أولية وصول مفتوح

LRCC: Generalizing Low-Rank Compression with Conditional Computation

Low-rank compression reduces the cost of pretrained language models by replacing linear transformations with low-rank factorizations. However, conventional methods use a fixed rank allocation during inference, assigning the same amount of compute regardless of the input token. We introduce Low-Rank Conditional Computat …

المؤلفون المشاركون