الباحثون

David Grangier

المنشورات 2

نسخة أولية وصول مفتوح

Lossy Compressive Text Autoencoders

Our work explores learning a compressed latent representation of text, at the intersection of data compression and representation learning. We propose an autoencoder architecture that performs residual downscaling and upscaling of hidden representations along the time axis, with a residual low-dimension discrete bottle …

نسخة أولية وصول مفتوح

KV-Kaizen: Learning Context-Adaptive Cache Compression Choices

As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This impacts LLM throughput negatively, since decoding is memory-bound and decode cost grows with cache size. Recent work alleviates this bottleneck by discarding the least re …

المؤلفون المشاركون