الباحثون

Fang Wu

المنشورات 3

نسخة أولية وصول مفتوح

Harness Evolution Hits a Ceiling: When Weight Training Should Begin

Yuan Tian, Bing Hu, Hao Wang وآخرون · 2026

Improving a long-horizon LLM agent means evolving the harness around a frozen model or training its weights. We let a self-evolving harness make the system stronger first, then cross seed and evolved harnesses with base and trained weights to learn which gains the trained model keeps and which still need the runtime. W …

نسخة أولية وصول مفتوح

Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight

Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction. For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful informat …

نسخة أولية وصول مفتوح

UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement

Fang Wu, Da Xing, Yanjie Huang وآخرون · 2026

Modern multimodal models bring generation and understanding into a single unified system, which enables them to provide and learn from their own feedback. Motivated by this unified capacity, we introduce UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback …

المؤلفون المشاركون