الباحثون

Wenhui Zhu

المنشورات 3

نسخة أولية وصول مفتوح

AvoKV-E: Payload-Aware KV Cache Eviction for Long Reasoning

Han Yu, Wenhui Zhu, Xiwen Chen وآخرون · 2026

Long-output reasoning shifts the KV-cache bottleneck from the fixed prompt to the generated trace. Existing reasoning-cache eviction methods largely treat cached entries as routing objects, estimating whether an old key will still be read, will recur, or can be replaced. This routing-only view overlooks two effects: lo …

نسخة أولية وصول مفتوح

Smaller Models, Better Rejects: Preference Distillation Scaling

Rui Cai, Wenhui Zhu, Xiwen Chen وآخرون · 2026

Preference distillation typically treats a teacher response as preferred and the student's own response as rejected. This assumes that self-generated failures are the most informative negatives and that rejects must come from a model at least as large as the student, making generation costly at scale. We find neither a …

المؤلفون المشاركون