الباحثون

Daize Dong

المنشورات 1

نسخة أولية وصول مفتوح

Broken Symmetry in BF16 Attention: Why FlashAttention Gradients Blow Up Late in Training

Junlin Chen, Daize Dong, Huanwei Di وآخرون · 2026

BF16 is now standard in large-scale pretraining, including in fused attention kernels such as FlashAttention, and these kernels are widely trusted. When we used FlashAttention-3 to pretrain a 450M-parameter transformer on 50B tokens, however, we ran into a problem: training was healthy for 25B tokens, then the gradient …

المؤلفون المشاركون