الباحثون

Yanggan Gu

المنشورات 5

نسخة أولية وصول مفتوح

TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning

Zhen Li, Shuai Zhang, Yanggan Gu وآخرون · 2026

Low-precision execution can substantially accelerate reinforcement learning (RL) for large language models, but discrepancies between learner and sampler execution can destabilize policy optimization. In this paper, we characterize the interaction between mismatch and the policy-gradient direction, distinguishing local …

نسخة أولية وصول مفتوح

InfiMed2: A Generalist Medical Multimodal Foundation Model from Contextual Evidence and Stability-Aware Supervision

Guanghao Zhu, Zeyu Liu, Zhitian Hou وآخرون · 2026

Recent medical multimodal models have benefited from larger corpora, broader modality coverage, and stronger reasoning-oriented training, yet effective data design across continued pretraining (CPT) and post-training remains challenging. Medical sources vary substantially in structure, granularity, and information dens …

نسخة أولية وصول مفتوح

SMAT: Simple and Efficient Merge-Aware Training

Yanggan Gu, Yuanyi Wang, Zhen Li وآخرون · 2026

Model merging integrates the capabilities of multiple experts without joint retraining, but standard expert training optimizes task loss alone and does not guarantee good performance after merging. Merge-aware training (MAT) aims to improve merged performance, but existing methods do not fully account for common mergin …

المؤلفون المشاركون