الباحثون

Xiao Li

المنشورات 9

نسخة أولية وصول مفتوح

LEAP: Making Privileged Geometry Supervision Effective for Visuomotor Learning

Han Fang, Yunpeng Jiang, Jianshu Hu وآخرون · 2026

Privileged 3D supervision uses additional geometric information during training to guide RGB-based visuomotor policy learning, without requiring geometric inputs at deployment. However, low reconstruction error does not ensure that visual representations capture geometry useful for control. We identify three limitation …

نسخة أولية وصول مفتوح

What Does Post-Training Change in Multilingual Reasoning?

Hongyang Li, Xiao Li, Caesar Wu وآخرون · 2026

Open-source reasoning models provide unequal access to reasoning capability across languages. When a model can solve a problem but cannot deliver a complete solution in the user's language, language becomes an access barrier rather than merely a source of performance variation. We audit Qwen3 checkpoints on competition …

نسخة أولية وصول مفتوح

Solving Without Stopping: On-Policy Distillation at Small Scale

Hongyang Li, Yiming Zhu, Xiao Li وآخرون · 2026

On-policy distillation, where a student learns from a stronger teacher's feedback on its own outputs, is a common way to pass reasoning to smaller models. We analyze what it transfers at small scale, distilling Qwen3-8B into Qwen3 4B, 1.7B and 0.6B students, in thinking mode (reason at length, then end the reasoning an …

نسخة أولية وصول مفتوح

Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models

Xiaoyu Luo, Tao Ren, Wenrui Yu وآخرون · 2026

The rapid capability gains of frontier language models are widely attributed to improved reasoning abilities, yet this cannot be verified as raw CoT traces in closed-source systems are hidden. By registering a simple custom tool through a standard API feature, we induce frontier models to externalize intermediate reaso …

نسخة أولية وصول مفتوح

Fewer Steps, Better Actions: Rethinking Flow-Matching Inference for VLA Policies

Zhipeng Tang, Xinda Chen, Weining Rao وآخرون · 2026

Vision-language-action (VLA) policies based on flow matching generate action chunks through repeated evaluations of an action expert. Increasing the number of integration steps raises inference cost, but does not necessarily improve closed-loop success. We propose Coda, which reallocates part of this integration budget …

المؤلفون المشاركون