الباحثون

Honglin Lin

المنشورات 2

نسخة أولية وصول مفتوح

MMVistaReason: Toward Open-Data and Post-Training Recipes for Multimodal Reasoning

Juekai Lin, Honglin Lin, Yuqian Yuan وآخرون · 2026

Open multimodal reasoning models have benefited from large-scale reasoning supervision, yet reliable post-training remains challenging due to uneven data quality, inefficient supervision construction, imbalanced difficulty, and cross-domain interference. We introduce MMVistaReason (MVR), an open-data post-training reci …

نسخة أولية وصول مفتوح

Teacher-Student Gaps Are Not Enough: Outcome-Guided On-Policy Distillation for Multi-Turn Autonomous Agents

Tong Zhang, Zhou Liu, Yihao Liu وآخرون · 2026

On-policy distillation (OPD) trains a student on its own trajectories with dense teacher supervision. Recent work on OPD for multi-turn autonomous agents often treats large teacher-student token-level distributional gaps as promising intervention points, linking larger gaps to a greater need for correction. Yet, our em …

المؤلفون المشاركون