الباحثون

Kai Chen

المنشورات 9

نسخة أولية وصول مفتوح

ResOPD: Tail Residualization for Sparse On-Policy Distillation

Penghui Yang, Long Xing, Xuanlang Dai وآخرون · 2026

On-policy distillation (OPD) trains a student model on self-generated trajectories, but transmitting dense teacher distributions across long reasoning traces creates prohibitive communication and memory bottlenecks. Practical systems therefore rely on sparse teacher interfaces, typically transmitting either the sampled …

نسخة أولية وصول مفتوح

Equivariant Visual-Tactile Diffusion Policy for Contact-Rich Manipulation

Lik Hang Kenny Wong, Yiyao Ma, Xiu-Shen Wei وآخرون · 2026

Imitation learning for contact-rich manipulation requires high-quality expert data that is expensive to obtain. This makes learning a sample-efficient policy a key issue. To address this, we propose VISTA, a workspace-level equivariant visuotactile diffusion policy for data-efficient contact-rich imitation learning. VI …

نسخة أولية وصول مفتوح

Gestalt: Large Multimodal Interplay Model

Zequn Yang, Yu Miao, Haotian Ni وآخرون · 2026

In this paper, we propose Gestalt, a new paradigm of large multimodal model built around multimodal interplay. Despite rapid advances, large multimodal models are reaching a bottleneck: existing approaches focus primarily on accommodating additional modalities while overlooking the distinct characteristics of each moda …

نسخة أولية وصول مفتوح

Can Agents Trust Their Skills? Uncovering Unsafe Chains of Trust in Skill-Based LLM Agents

Yan Wang, Zhihao Zhang, Ke Chen وآخرون · 2026

LLM agents increasingly rely on installable skills, which are packages of instructions, code, and resources that equip them with task-specific capabilities and, once installed, can be automatically invoked across subsequent user tasks. This creates a chain of trust in which users delegate authority to agents, while age …

نسخة أولية وصول مفتوح

Cooperative Multi-Agent Vision-Language-Action Models via Reinforced Fine Tuning

Ruixiao Xu, Wong Lik Hang Kenny, Zhiqian Liu وآخرون · 2026

We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLAs are pretrained on large-scale single-agent data and therefore lack the fine-grained coordination skills required for inter-robot collaboration. Supervised fine-tuning (S …

نسخة أولية وصول مفتوح

SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving

Zhilong Ge, Yuting Shao, Yutao Yang وآخرون · 2026

Human-written agent skills encode rich workflows for real-world problem solving, but are typically used as external inference-time instructions rather than internalized as reusable model capabilities. We introduce \texttt{SkillGym}, a framework that transforms these skills into executable, verifiable training environme …

نسخة أولية وصول مفتوح

GTR: Gated Token Recurrence for Efficient Dense Prediction

Zhe Feng, Longfei Liu, Wei Liu وآخرون · 2026

Self-attention-based vision backbones perform well on dense prediction, but the quadratic computational cost of global softmax attention limits their efficiency as image resolution increases. We introduce Gated Token Recurrence (GTR), a softmax-free recurrent vision backbone that combines gated linear attention, altern …

نسخة أولية وصول مفتوح

ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

Shijie Lian, Bin Yu, Zhaolong Shen وآخرون · 2026

Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet sma …

المؤلفون المشاركون