الملخص

Data selection is already a central bottleneck in large-language-model training, where web-scale corpora are noisy and token budgets are finite. In continual pre-training (CPT), it becomes a forgetting-control problem: a poorly chosen target-domain corpus can overwrite capabilities encoded in the pretrained checkpoint. Existing CPT practice either scores candidates with parameter-agnostic scalars such as perplexity, or mitigates forgetting by spending many extra general-domain replay tokens. Neither strategy directly asks how training on a candidate will move the model parameters. We show that loss-based selection causes the post-CPT Fisher diagonal to drift downward on exactly the high-Fisher coordinates the pretrained model had committed to, while leaving low-Fisher coordinates largely untouched. This asymmetry exposes a parameter-space mechanism for catastrophic forgetting. Motivated by this observation, we propose a Fisher-aware CPT selector that decomposes each candidate's gradient into an anchor component, which measures perturbation along committed parameter directions, and a frontier component, which measures update capacity in unconstrained low-Fisher subspaces. We aggregate these signals with a log-determinant submodular objective and optimize it in a single pass using a scalable streaming data selection pipeline. On TinyLlama-1.1B and Llama-3.1-8B CPT over medical data, our selector improves target-domain quality while bounding forgetting on held-out pretraining benchmarks. Most importantly, it is substantially more token-efficient than forgetting-aware replay. 1B selected tokens already outperform the replay strategy trained with 10B tokens on both adaptation and forgetting, giving a 10x token-efficiency advantage.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Zhao, Z., Liu, G., Lan, Z., & Yan, Y. (2026). Fisher-Guided Submodular Data Selection for Continual Pre-Training of Large Language Models. https://omanscience.com/ar/articles/fisher-guided-submodular-data-selection-for-continual-pre-training-of-large-language-models

MLA 9

Zhao, Zhenghao, et al. "Fisher-Guided Submodular Data Selection for Continual Pre-Training of Large Language Models." https://omanscience.com/ar/articles/fisher-guided-submodular-data-selection-for-continual-pre-training-of-large-language-models.

شيكاغو (المؤلف–التاريخ)

Zhao, Zhenghao, Gaowen Liu, Zhiling Lan, and Yan Yan. 2026. "Fisher-Guided Submodular Data Selection for Continual Pre-Training of Large Language Models." https://omanscience.com/ar/articles/fisher-guided-submodular-data-selection-for-continual-pre-training-of-large-language-models.

هارفارد

Zhao, Z., Liu, G., Lan, Z. and Yan, Y. (2026) 'Fisher-Guided Submodular Data Selection for Continual Pre-Training of Large Language Models', Available at: https://omanscience.com/ar/articles/fisher-guided-submodular-data-selection-for-continual-pre-training-of-large-language-models.

فانكوفر

Zhao Z, Liu G, Lan Z, Yan Y. Fisher-Guided Submodular Data Selection for Continual Pre-Training of Large Language Models. https://omanscience.com/ar/articles/fisher-guided-submodular-data-selection-for-continual-pre-training-of-large-language-models

IEEE

Z. Zhao, G. Liu, Z. Lan, and Y. Yan, "Fisher-Guided Submodular Data Selection for Continual Pre-Training of Large Language Models," https://omanscience.com/ar/articles/fisher-guided-submodular-data-selection-for-continual-pre-training-of-large-language-models.