الباحثون

Sirui Han

المنشورات 5

نسخة أولية وصول مفتوح

HarnessSQL: Harness-Native Training for SQL Agents in Realistic Database Environments

Haolin Yang, Jipeng Zhang, Jian Xie وآخرون · 2026

Text-to-SQL models are commonly trained to map questions directly to static queries, whereas real-world database agents operate through stateful, multi-turn interaction with live databases -- inspecting schemas, executing probe queries, diagnosing errors, and revising hypotheses. This creates a critical train-deploy mi …

نسخة أولية وصول مفتوح

Token-World: World Modeling in Vision-Language Model Token Space for Robot Manipulation

Chuyao Fu, Xiaowei Chi, Yuhan Rui وآخرون · 2026

A common approach to world-model simulation for vision-language-action (VLA) systems is to predict future RGB observations and then re-encode them into policy inputs, introducing an indirect interface between simulation and downstream policy execution. We instead investigate whether world dynamics can be modeled in a c …

نسخة أولية وصول مفتوح

CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding

Yulong Liu, Xiaotian Han, Junyuan Shang وآخرون · 2026

Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, making long-video understanding expensive for the vision encoder and the language model. Existing methods often compress visual tokens after dense encoding, creating a mism …

نسخة أولية وصول مفتوح

DexRoam: Learning Mobile Bimanual Dexterous Manipulation from Egocentric Whole-Body Human Demonstrations

Rui Zhou, Yibo Yuan, Junkai Zhao وآخرون · 2026

Mobile bimanual dexterous manipulation requires continuous coordination of locomotion, whole-body motion, and finger-level dexterity within a single trajectory, creating a severe robot demonstration bottleneck. Egocentric human demonstrations offer a scalable alternative, but prior approaches ease the transfer by simpl …

المؤلفون المشاركون