Authors

Han Yu

Publications 6

Preprint Open access

AvoKV-E: Payload-Aware KV Cache Eviction for Long Reasoning

Han Yu, Wenhui Zhu, Xiwen Chen et al. · 2026

Long-output reasoning shifts the KV-cache bottleneck from the fixed prompt to the generated trace. Existing reasoning-cache eviction methods largely treat cached entries as routing objects, estimating whether an old key will still be read, will recur, or can be replaced. This routing-only view overlooks two effects: lo …

Preprint Open access

Smaller Models, Better Rejects: Preference Distillation Scaling

Rui Cai, Wenhui Zhu, Xiwen Chen et al. · 2026

Preference distillation typically treats a teacher response as preferred and the student's own response as rejected. This assumes that self-generated failures are the most informative negatives and that rejects must come from a model at least as large as the student, making generation costly at scale. We find neither a …

Preprint Open access

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Peng Xia, Rujun Han, Zifeng Wang et al. · 2026

An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically e …

Preprint Open access

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

DeepSeek-AI, Anyi Xu, B. Li et al. · 2026

The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Togeth …

Co-authors