الباحثون

بينغسيانغ لي

المنشورات 3

نسخة أولية وصول مفتوح

Scaling to Tens of Thousands of Test-Time Iterations with Loop-Native Attention Residuals

In this paper, we argue that looped Transformers need their own residual connections to prevent performance degradation as the number of iterations grows. We observe that increasing loop iterations can reduce reasoning accuracy: noisy state updates overwrite correct intermediate deductions and even undo completed solut …

نسخة أولية وصول مفتوح

Plan Canvas: Fixed Reasoning Regions for Continuous Language Flows

Continuous language flows generate text by denoising all positions of a target canvas together. The natural way to add reasoning to such a model is to write a trace ahead of the answer, but the trace length changes from question to question. The answer start is therefore unknown during denoising, and the model has to d …

نسخة أولية وصول مفتوح

Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts

Looped Transformers introduce recurrent depth as a new scaling axis for LLMs: by repeatedly applying shared Transformer blocks, they increase effective depth without increasing parameter count. However, the benefits of looping remain unclear for large MoE LLMs under FLOPs-matched comparisons. The main reason is that th …

المؤلفون المشاركون