الباحثون

Zhenhao Zhang

المنشورات 2

نسخة أولية وصول مفتوح

ReTaCo: Residual-Target Control for On-Policy Distillation

Zixiang Ni, Zhuo Hu, Renjie Cao وآخرون · 2026

On-policy distillation (OPD) trains a student on its own generated prefixes with token-level teacher feedback, but transmitting or storing the teacher's full-vocabulary distribution at every token is costly. Entropy-aware OPD (EOPD) adds forward supervision to reverse KL to help the student recover plausible tokens it …

نسخة أولية وصول مفتوح

CaptchaArena: A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs

Zhenhao Zhang, Zhaoyu Fan, Haohan Ying وآخرون · 2026

Interactive CAPTCHAs remain challenging for computer-use agents, while existing datasets face trade-offs among type coverage, interaction fidelity, and trajectory supervision. To address these gaps, we present CaptchaArena, the first large-scale, fine-grained training dataset for interactive CAPTCHA solving. It contain …

المؤلفون المشاركون