الباحثون

Yixuan Chen

المنشورات 3

نسخة أولية وصول مفتوح

The Functional Structure of Post-Compression Recovery in Low-Rank LLMs

Zishan Shao, Liang Tian, Georgiy Zemlevskiy وآخرون · 2026

Different low-rank compression methods can produce compressed LLMs that respond differently to the same post-compression recovery procedure, and relative advantages observed between methods at the compression endpoint may shrink, grow, or even reverse after recovery. We ask whether this recovery heterogeneity reflects …

نسخة أولية وصول مفتوح

Joint Branch-Space Transform Coding for Diffusion Activation Quantization with Classifier-Free Guidance

Mingrun Jiang, Yuejia Liu, Zishan Shao وآخرون · 2026

Post-training quantization for diffusion models increasingly exploits timestep, feature, and layer structure. While recent work has begun incorporating CFG structure into diffusion quantization, activation quantization still operates independently across conditional and unconditional coordinates, leaving cross-activati …

نسخة أولية وصول مفتوح

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

DeepSeek-AI, Anyi Xu, B. Li وآخرون · 2026

The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Togeth …

المؤلفون المشاركون