الباحثون

Feida Zhu

المنشورات 2

نسخة أولية وصول مفتوح

ProCredit: From Outcome Rewards to Progress Credit in Agentic Reinforcement Learning

Ming Ma, Yi Zhu, Yiran Zhong وآخرون · 2026

Long-horizon agentic tasks require an agent to modify an environment through a sequence of tool calls, with success determined by the final state. The standard recipe assigns a single outcome reward at the end and compares trajectories sampled for the same task. As a result, a group with no successful trajectory yields …

نسخة أولية وصول مفتوح

Block-Sparse Attention with Semantic-Geometric Decoupled Routing

Xinwei Long, Weigao Sun, Weibo Gao وآخرون · 2026

Long-context inference has become a defining capability of large language models, but exact dense attention remains costly due to its quadratic scaling with sequence length. Block-sparse attention offers a hardware-friendly alternative by routing each query block to a small set of relevant key blocks, yet accurate trai …

المؤلفون المشاركون