الملخص
Test-time training (TTT) allows a model to improve its predictions at inference time by updating weights after every observed token. However, sequential gra- dient writes make parallel training difficult. We observe that, given layer inputs and activation gradients (costates), online gradient descent admits exact parallel scans for both forward evaluation and reverse backpropagation. GradLev lever- ages this duality: a causal auxiliary network predicts costates across all tokens in parallel; associative scans compute the adapted weights and forward activations and propagate gradients backward; and the resulting gradient targets supervise the predictor via a consistency loss. Exact consistency guarantees exact recovery of the sequential online learner. At deployment, the auxiliary predictor is discarded, and the model updates natively via token-by-token forward and backward passes.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Liu, B., & Liu, Q. (2026). GradLev: Token-Parallel Test-Time Training Via Costate Prediction. https://omanscience.com/ar/articles/gradlev-token-parallel-test-time-training-via-costate-prediction
MLA 9
Liu, Bo, and Qiang Liu. "GradLev: Token-Parallel Test-Time Training Via Costate Prediction." https://omanscience.com/ar/articles/gradlev-token-parallel-test-time-training-via-costate-prediction.
شيكاغو (المؤلف–التاريخ)
Liu, Bo, and Qiang Liu. 2026. "GradLev: Token-Parallel Test-Time Training Via Costate Prediction." https://omanscience.com/ar/articles/gradlev-token-parallel-test-time-training-via-costate-prediction.
هارفارد
Liu, B. and Liu, Q. (2026) 'GradLev: Token-Parallel Test-Time Training Via Costate Prediction', Available at: https://omanscience.com/ar/articles/gradlev-token-parallel-test-time-training-via-costate-prediction.
فانكوفر
Liu B, Liu Q. GradLev: Token-Parallel Test-Time Training Via Costate Prediction. https://omanscience.com/ar/articles/gradlev-token-parallel-test-time-training-via-costate-prediction
IEEE
B. Liu, and Q. Liu, "GradLev: Token-Parallel Test-Time Training Via Costate Prediction," https://omanscience.com/ar/articles/gradlev-token-parallel-test-time-training-via-costate-prediction.