الباحثون

Zhirui Zhang

المنشورات 3

نسخة أولية وصول مفتوح

Act First, Reason Later: Accelerating On-Policy Distillation for Multi-Turn Agents via Reference-Conditioned Inverse Dynamics

Zubin Zheng, Jiahao Wu, Shaofeng Zhang وآخرون · 2026

On-policy distillation (OPD) trains multi-turn language agents with dense teacher supervision on student-generated responses. However, standard think-then-act rollouts require lengthy reasoning before each short action, delaying environment transitions and experience collection. Generating actions directly reduces this …

نسخة أولية وصول مفتوح

ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

Shijie Lian, Bin Yu, Zhaolong Shen وآخرون · 2026

Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet sma …

المؤلفون المشاركون