الباحثون

Bernard Nguyen

المنشورات 1

نسخة أولية وصول مفتوح

NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agentic RL at Trillion-Parameter Scale

Songlin Jiang, Zhiyu Li, Terry Kong وآخرون · 2026

Agentic reinforcement learning (RL) disaggregates training from rollout, so each policy update must reach the rollout clusters before the next batch. Transferring a full 1T checkpoint for such weight synchronization (refit) takes 87.5 min between two AWS regions. Measurements of BF16 training show that about 1% of weig …

المؤلفون المشاركون