الباحثون

Jiahao Lu

المنشورات 4

نسخة أولية وصول مفتوح

Agentic-TTT: Training test-time policy for test-time training

Test-time training (TTT) adapts an LLM's parameters using signals derived from test inputs, and can make striking improvements in pre-specified settings such as IMO competitions or designated open problems. By turning deployment experience into parameter updates, TTT provides a direct mechanism for model-level self-imp …

نسخة أولية وصول مفتوح

MARCO: Multi-Round Agentic Reinforcement for Conditional Molecular Optimization

Shicheng Fang, Yuxin Wang, Zhuo Yang وآخرون · 2026

Molecular optimization is inherently iterative: a candidate is proposed, evaluated against several objectives, and revised while preserving a relationship to the source molecule. Most instruction-following models instead emit one edited molecule, forcing validity, property improvement, and similarity control into a sin …

نسخة أولية وصول مفتوح

ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients

Shicheng Fang, Yiwen Zhao, Wenbo Tian وآخرون · 2026

Multi-reward policy optimization requires a joint update that reflects both the learning signals and the intended relationships among objectives. We introduce Objective-wise Reconciled Policy Gradient (ORPG), which constructs a separate clipped policy objective for each reward and reconciles the resulting gradients int …

المؤلفون المشاركون