الباحثون

Xipeng Qiu

المنشورات 7

نسخة أولية وصول مفتوح

XTurnix: Large-Scale Self-Supervised Turn Control through Two-State Binary Decisions

Zhanxun Liu, Yifan Duan, Hengtao Wu وآخرون · 2026

General turn-taking behavior in real-time dialogue systems requires deciding whether to keep listening or start responding while listening, and whether to continue or stop while speaking. Existing turn detectors use heterogeneous, task-specific label spaces and are often trained on limited annotations or evaluated on i …

نسخة أولية وصول مفتوح

MARCO: Multi-Round Agentic Reinforcement for Conditional Molecular Optimization

Shicheng Fang, Yuxin Wang, Zhuo Yang وآخرون · 2026

Molecular optimization is inherently iterative: a candidate is proposed, evaluated against several objectives, and revised while preserving a relationship to the source molecule. Most instruction-following models instead emit one edited molecule, forcing validity, property improvement, and similarity control into a sin …

نسخة أولية وصول مفتوح

WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents

Bo Mao, Hang He, Linting Wang وآخرون · 2026

Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scaling environments in isolation, however, does not guarantee com …

نسخة أولية وصول مفتوح

SPIDER: Multi-Layer Semantic Token Pruning and Adaptive Sub-Layer Skipping in Multimodal Large Language Models

Tianxiang Chen, Zhentao Tan, Zi Ye وآخرون · 2026

Multimodal Large Language Models face significant efficiency challenges that stem from two distinct yet coupled sources: data redundancy and computational redundancy. While most methods focus on data redundancy by pruning visual tokens from the output of the visual encoder or computing redundancy in LLM decoders using …

نسخة أولية وصول مفتوح

ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients

Shicheng Fang, Yiwen Zhao, Wenbo Tian وآخرون · 2026

Multi-reward policy optimization requires a joint update that reflects both the learning signals and the intended relationships among objectives. We introduce Objective-wise Reconciled Policy Gradient (ORPG), which constructs a separate clipped policy objective for each reward and reconciles the resulting gradients int …

المؤلفون المشاركون