الباحثون

Yuxin Wang

المنشورات 5

نسخة أولية وصول مفتوح

Efficient Best-of-N policy evaluation for inference-time alignment

Best-of-N (BoN) is a common inference-time alignment method that selects the highest-scoring response among N samples from a reference model. Evaluating BoN policies from logged data is challenging under sample-only access because standard off-policy estimators require density ratios that depend on unavailable response …

نسخة أولية وصول مفتوح

Prediction-powered Neural Architecture Search

Pascal Janetzky, Yuxin Wang, Michael Klar وآخرون · 2026

Evaluating candidate architectures in neural architecture search (NAS) faces an inherent trade-off: on the one hand, reliable performance labels are limited because training and evaluating architectures is costly; on the other hand, zero-cost proxies (ZCPs) are cheap to compute at large scale but can be noisy. Yet, how …

نسخة أولية وصول مفتوح

MARCO: Multi-Round Agentic Reinforcement for Conditional Molecular Optimization

Shicheng Fang, Yuxin Wang, Zhuo Yang وآخرون · 2026

Molecular optimization is inherently iterative: a candidate is proposed, evaluated against several objectives, and revised while preserving a relationship to the source molecule. Most instruction-following models instead emit one edited molecule, forcing validity, property improvement, and similarity control into a sin …

نسخة أولية وصول مفتوح

ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients

Shicheng Fang, Yiwen Zhao, Wenbo Tian وآخرون · 2026

Multi-reward policy optimization requires a joint update that reflects both the learning signals and the intended relationships among objectives. We introduce Objective-wise Reconciled Policy Gradient (ORPG), which constructs a separate clipped policy objective for each reward and reconciles the resulting gradients int …

المؤلفون المشاركون