الباحثون

Yige Wang

المنشورات 3

نسخة أولية وصول مفتوح

CERO: Where and When to Allocate Rollouts for RL Post-Training

Yiming Zong, Yige Wang, Xing Hu وآخرون · 2026

Adaptive rollout methods for group-relative reinforcement learning typically allocate a fixed per-update budget across prompts. We instead study how to coordinate a finite rollout budget over the entire training horizon. We formulate this problem using a concave surrogate utility of cumulative prompt exposure and intro …

نسخة أولية وصول مفتوح

How Can Recommendation Feedback Evolve Agent Memory?

Shanwen Mao, Mingming Li, Hao Zhang وآخرون · 2026

Content-generation agents continuously receive impressions, clicks, conversions, and negative feedback from recommendation systems, providing real-world outcome signals for memory evolution. However, these signals are delayed and noisy, confounded by audience composition, placement, and recommendation policies, and may …

المؤلفون المشاركون