الباحثون

Linchao Zhu

المنشورات 3

نسخة أولية وصول مفتوح

MOBA-VL: Event-Localized Multi-Turn Reinforcement Learning for Real-Time MOBA Commentary

Shengyun Zhong, Xinkang Zhao, Ziyuan Chu وآخرون · 2026

Real-time commentary for Multiplayer Online Battle Arena (MOBA) esports requires a vision-language model (VLM) to narrate a live match second by second, both fluently and accurately. Existing streaming VLMs sound natural but often miss key events such as kills and objectives. To address this limitation, we use game tel …

نسخة أولية وصول مفتوح

RAVEL: Asynchronous Rolling Inference for Flow-Based Vision-Language-Action Models

Yuhan Chen, Ke Yu, Pengfei Liu وآخرون · 2026

Flow-based vision-language-action (VLA) models are highly effective for generalist robot manipulation, yet their reliance on computationally expensive VLM encoding and multi-step iterative action generation imposes a significant latency bottleneck. The resulting inference latency makes it difficult for robots to respon …

نسخة أولية وصول مفتوح

CaptchaArena: A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs

Zhenhao Zhang, Zhaoyu Fan, Haohan Ying وآخرون · 2026

Interactive CAPTCHAs remain challenging for computer-use agents, while existing datasets face trade-offs among type coverage, interaction fidelity, and trajectory supervision. To address these gaps, we present CaptchaArena, the first large-scale, fine-grained training dataset for interactive CAPTCHA solving. It contain …

المؤلفون المشاركون