الباحثون

Hao Yu

المنشورات 6

نسخة أولية وصول مفتوح

TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models

Xin Wang, Hao Yu, Zhengyang Zhuge وآخرون · 2026

Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily optimize quantization ac …

نسخة أولية وصول مفتوح

Shallow Queries, Mature Values: Depth-Asynchronous Self-Speculation for Looped Transformers

Guanghao Li, Zihan Su, Hao Yu وآخرون · 2026

Looped Transformers reuse a shared block across recurrent depths, making autoregressive decoding expensive because every generated token requires many sequential recurrent passes. Self-speculative decoders reduce this cost by drafting at an early depth and verifying at full depth, but typically bind draft computation t …

المؤلفون المشاركون