Authors

Yang Zhang

Publications 14

Preprint Open access

Learning to Accumulate Knowledge with Mutual Information

Yuyang Zhao, Lizi Liao, Leyang Shen et al. · 2026

Large language model (LLM) agents can improve their performance by reusing knowledge distilled from past interactions. However, curating new experiences into a knowledge bank that becomes more useful as it grows remains challenging. Effective knowledge accumulation should limit redundant overlap among entries and ensur …

Preprint Open access

SHarP: Saliency-based Pruning of Agent Harnesses

Xinyi Gao, Qiucheng Wu, Kaizhi Qian et al. · 2026

Agent harnesses are systems that coordinate model calls, tool use, and task execution to help large language models complete complex tasks. To meet task requirements and address failures, these systems are often iteratively refined by amending and patching their instructions, tools, and workflows, continuously increasi …

Preprint Open access

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

DeepSeek-AI, Anyi Xu, B. Li et al. · 2026

The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Togeth …

Co-authors