الباحثون

Zhiling Lan

المنشورات 1

نسخة أولية وصول مفتوح

Fisher-Guided Submodular Data Selection for Continual Pre-Training of Large Language Models

Zhenghao Zhao, Gaowen Liu, Zhiling Lan وآخرون · 2026

Data selection is already a central bottleneck in large-language-model training, where web-scale corpora are noisy and token budgets are finite. In continual pre-training (CPT), it becomes a forgetting-control problem: a poorly chosen target-domain corpus can overwrite capabilities encoded in the pretrained checkpoint. …

المؤلفون المشاركون