الملخص
A central challenge in continual learning is to acquire new knowledge without forgetting what the model has already learned. This challenge appears in language model training when training data comes from various domain-, user-, or task-specific distributions that are encountered unevenly over time. In such settings, successive minibatches are temporally clustered by distribution instead of being sampled i.i.d. from the overall data mixture. Training on temporally clustered data induces a stability-plasticity tradeoff. Adapting the model to the active distribution can improve the model on the active distribution but may lead to a performance degradation on data it previously trained on. We find that this tradeoff intensifies with longer exposure to the same distribution. We therefore ask if the parameter displacement produced by such a sequence (the task vector) should be fully retained or applied partially. We compare applying the full displacement ($λ=1$) with partial integration, which scales the task vector by $λ$ before applying it to the continuing model and scales the optimizer state by the same coefficient. Across continual pretraining, pretraining from random initialization, supervised post-training, and reinforcement post-training, we find that intermediate values of $λ$ often improve average continuing-model performance relative to full integration, particularly after longer same-distribution sequences. In continual-pretraining experiments with both controlled streams and naturally defined adaptation sequences, task-vector scaling outperforms full integration at the matched learning rate, showing that its benefits are not reproduced by learning-rate scaling alone.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Baumann, A., Hübotter, J., Akata, Z., & Krause, A. (2026). Task Vector Descent: Learning from Non-IID Batches. https://omanscience.com/ar/articles/task-vector-descent-learning-from-non-iid-batches
MLA 9
Baumann, Anton, et al. "Task Vector Descent: Learning from Non-IID Batches." https://omanscience.com/ar/articles/task-vector-descent-learning-from-non-iid-batches.
شيكاغو (المؤلف–التاريخ)
Baumann, Anton, Jonas Hübotter, Zeynep Akata, and Andreas Krause. 2026. "Task Vector Descent: Learning from Non-IID Batches." https://omanscience.com/ar/articles/task-vector-descent-learning-from-non-iid-batches.
هارفارد
Baumann, A., Hübotter, J., Akata, Z. and Krause, A. (2026) 'Task Vector Descent: Learning from Non-IID Batches', Available at: https://omanscience.com/ar/articles/task-vector-descent-learning-from-non-iid-batches.
فانكوفر
Baumann A, Hübotter J, Akata Z, Krause A. Task Vector Descent: Learning from Non-IID Batches. https://omanscience.com/ar/articles/task-vector-descent-learning-from-non-iid-batches
IEEE
A. Baumann, J. Hübotter, Z. Akata, and A. Krause, "Task Vector Descent: Learning from Non-IID Batches," https://omanscience.com/ar/articles/task-vector-descent-learning-from-non-iid-batches.