الباحثون

Dezhi Ran

المنشورات 2

نسخة أولية وصول مفتوح

Backward-State Policy Is Part of the Learning Algorithm

Shuxiao Xie, Shuyang Xie, Dezhi Ran وآخرون · 2026

Low-precision training rounds tensors that the backward pass reads again, often for several gradients; each use can read the forward's rounded value, the original, or a new random rounding. This backward-state policy looks like a memory and precision detail, settled by copy accuracy and final loss. We argue that it is …

نسخة أولية وصول مفتوح

Beyond Accuracy: Prefix-Invariant Realizations of Low-Precision Fast Matrix Multiplication

Shuxiao Xie, Shuyang Xie, Yuan Cao وآخرون · 2026

Fast matrix multiplication saves multiplications through exact cancellation, but rounding sums that mix token rows can leave contributions from later tokens in earlier language model outputs. This threatens prefix invariance, which multiple-choice likelihood scoring relies on: a scored likelihood must depend only on it …

المؤلفون المشاركون