الباحثون

Haotian Xu

المنشورات 2

نسخة أولية وصول مفتوح

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

DeepSeek-AI, Anyi Xu, B. Li وآخرون · 2026

The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Togeth …

نسخة أولية وصول مفتوح

Bellman Policy Optimization

Zhuoqing Song, Haotian Xu, Xikun Zhang وآخرون · 2026

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models (LLMs). We introduce Bellman Policy Optimization (BPO), a critic-free method derived from Policy Mirror Descent (PMD). For autoregressive generation with terminal rewards, BPO uses the Bellman equations to …

المؤلفون المشاركون