الباحثون

Dan Alistarh

المنشورات 3

نسخة أولية وصول مفتوح

Q-PACE: Dynamic Precision Allocation for Quantization-Aware Training

Quantization-aware training (QAT) leverages lower-precision arithmetic to reduce the cost of LLM deployment, but aggressive quantization degrades final model performance. A common remedy is mixed-precision training, in which high precision is assigned to some of the layers to maintain performance while keeping the cost …

نسخة أولية وصول مفتوح

WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms

Jiale Chen, Vage Egiazarian, Eldar Kurtić وآخرون · 2026

KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-order statistics of both factors in a ma …

المؤلفون المشاركون