الباحثون

Hongyi Fu

المنشورات 2

نسخة أولية وصول مفتوح

Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL

Youling Huang, Tiankuo Xu, Jiaji Liu وآخرون · 2026

Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-stage policies often fail to solve sampled tasks, leaving little useful reward signal for learning. To mitigate this problem, we use on-policy distillation (OPD) to provid …

نسخة أولية وصول مفتوح

GSM: Efficient Language Modeling with Shared Global State

Yunao Zheng, Bin Wen, Xiaojie Wang وآخرون · 2026

Efficient language models must reduce not only the cost of individual accesses to past context but also the overhead of repeatedly selecting and processing historical information across layers. We introduce the Global State Model (GSM), a causal encoder--decoder architecture that concentrates the selection and aggregat …

المؤلفون المشاركون