الملخص
Supervised fine-tuning (SFT) adapts pretrained large language models (LLMs) to downstream tasks, but the required concepts can receive substantially different levels of pretrained support. Frequent concepts are more likely to be well learned, whereas rare concepts may remain weakly represented. We introduce a novel notion named prior barrier to quantify how strongly the pretrained model supports competing concepts over the target concept. We observe that prior barriers follow a long-tail distribution, placing head and tail concepts at different starting points for SFT: head concepts face lower prior barriers, whereas tail concepts require additional instructions to overcome their higher prior barriers. Our theoretical analysis further derives a predictive risk bound for SFT under long-tail prior barriers, explicitly characterizing how the prior barrier and accumulated SFT evidence jointly determine predictive performance. Motivated by this prior barrier-dependent demand, we propose PASS, an adaptive SFT instruction selection method that constructs reference-derived concepts and estimates the distinguishing evidence provided by each instruction, and adaptively allocates the selection budget toward concepts that remain insufficiently covered under the current selection. In this way, PASS jointly considers which instructions can provide useful evidence and where additional supervision is needed under a limited budget. Experiments show that our method consistently outperforms seven state-of-the-art instruction selection methods on four backbone-budget settings. An ablation study further shows that PASS's adaptive allocation consistently improves over uniform allocation.
الكلمات المفتاحية
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Wang, H., Xu, J., Zhan, W., Zeng, T., Fu, D., Li, H., Roy, S., Ramakrishnan, N., North, C., Kang, J., Yan, Y., & Zhou, D. (2026). Overcoming Prior Barriers: Supervised Fine-Tuning under Long-Tail Distribution. https://omanscience.com/ar/articles/overcoming-prior-barriers-supervised-fine-tuning-under-long-tail-distribution
MLA 9
Wang, Haohui, et al. "Overcoming Prior Barriers: Supervised Fine-Tuning under Long-Tail Distribution." https://omanscience.com/ar/articles/overcoming-prior-barriers-supervised-fine-tuning-under-long-tail-distribution.
شيكاغو (المؤلف–التاريخ)
Wang, Haohui, Jiahao Xu, Wangzhi Zhan, Tong Zeng, Dongqi Fu, Hong Li, Swastik Roy, Naren Ramakrishnan, Chris North, Jian Kang, Yujun Yan, and Dawei Zhou. 2026. "Overcoming Prior Barriers: Supervised Fine-Tuning under Long-Tail Distribution." https://omanscience.com/ar/articles/overcoming-prior-barriers-supervised-fine-tuning-under-long-tail-distribution.
هارفارد
Wang, H., Xu, J., Zhan, W., Zeng, T., Fu, D., Li, H., Roy, S., Ramakrishnan, N., North, C., Kang, J., Yan, Y. and Zhou, D. (2026) 'Overcoming Prior Barriers: Supervised Fine-Tuning under Long-Tail Distribution', Available at: https://omanscience.com/ar/articles/overcoming-prior-barriers-supervised-fine-tuning-under-long-tail-distribution.
فانكوفر
Wang H, Xu J, Zhan W, Zeng T, Fu D, Li H, et al. Overcoming Prior Barriers: Supervised Fine-Tuning under Long-Tail Distribution. https://omanscience.com/ar/articles/overcoming-prior-barriers-supervised-fine-tuning-under-long-tail-distribution
IEEE
H. Wang, J. Xu, W. Zhan, T. Zeng, D. Fu, H. Li, S. Roy, N. Ramakrishnan, C. North, J. Kang, Y. Yan, and D. Zhou, "Overcoming Prior Barriers: Supervised Fine-Tuning under Long-Tail Distribution," https://omanscience.com/ar/articles/overcoming-prior-barriers-supervised-fine-tuning-under-long-tail-distribution.