الباحثون

Xiao Zhou

المنشورات 2

نسخة أولية وصول مفتوح

On the Efficiency-Safety Dilemma in Large Reasoning Models

Yifei Yang, Zouying Cao, Xingrui Wang وآخرون · 2026

Large reasoning models (LRMs) incur high inference costs, often mitigated by efficiency techniques like quantization and pruning. However, the impact of these techniques on model adversarial robustness remains largely unexplored. This study provides the first comprehensive analysis of the interplay between efficiency, …

نسخة أولية وصول مفتوح

GrowMTP: Can RL Grow Its Own Draft Head?

Minghua He, Lingzhe Zhang, Yuan Liu وآخرون · 2026

Reinforcement learning (RL) post-training drives the frontier capabilities of large language models, with its wall-clock dominated by autoregressive rollout generation. Speculative decoding is an established remedy for this bottleneck, but existing draft heads must be pretrained or warmed up before RL, introducing subs …

المؤلفون المشاركون