الباحثون

Haoxiang Xu

المنشورات 2

نسخة أولية وصول مفتوح

Dense Is Not Enough: Hierarchical Supervision Allocation for Long-Horizon On-Policy Distillation

Yuhao Sun, Binrui Wu, Zhuoer Xu وآخرون · 2026

On-policy distillation (OPD) transfers the capabilities of a large language model to a smaller student by providing teacher supervision on the student's own rollouts. In long-horizon agentic tasks, however, uniform token-level matching can allocate supervision poorly: a large local discrepancy need not improve future b …

نسخة أولية وصول مفتوح

Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction

Lujia Bao, Qian Chen, Luyao Cheng وآخرون · 2026

Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these requirements together through Think, Act, and Speak and Coordinate. Think combines Core-Cocktail supervised fine-tuning with Multimodality and Multi-Teacher On-Policy Dis …

المؤلفون المشاركون