Abstract
Adapting pretrained models to downstream tasks with limited data has become a central paradigm in modern deep learning. Yet, despite its widespread practical success, how fine-tuning leverages information from pretraining remains poorly understood theoretically. We study fine-tuning from pretrained weights through the lens of sparse linear regression and two-layer diagonal linear networks. In our setting, pretraining provides information through the support (and signs) of the initialization predictor, which may contain coordinates relevant to the downstream task. We show how pretrained information reshapes the implicit bias and training dynamics, and can thereby reduce the sample complexity of recovering the target parameters and support. In particular, for a clean initialization with correctly inherited signs, we show that the required sample size is comparable to that of a weighted Lasso estimator that explicitly exploits the pretrained support through a suitably chosen regularizer. Our results thus show how information encoded in pretrained weights can be implicitly exploited by gradient-based fine-tuning, reducing the amount of data needed to recover a downstream task.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Declèves, A., Boursier, E., & Flammarion, N. (2026). Statistical Benefits of Fine-Tuning from Pretrained Initialization in Diagonal Linear Networks. https://omanscience.com/en/articles/statistical-benefits-of-fine-tuning-from-pretrained-initialization-in-diagonal-linear-networks
MLA 9
Declèves, Alexandre, et al. "Statistical Benefits of Fine-Tuning from Pretrained Initialization in Diagonal Linear Networks." https://omanscience.com/en/articles/statistical-benefits-of-fine-tuning-from-pretrained-initialization-in-diagonal-linear-networks.
Chicago (author–date)
Declèves, Alexandre, Etienne Boursier, and Nicolas Flammarion. 2026. "Statistical Benefits of Fine-Tuning from Pretrained Initialization in Diagonal Linear Networks." https://omanscience.com/en/articles/statistical-benefits-of-fine-tuning-from-pretrained-initialization-in-diagonal-linear-networks.
Harvard
Declèves, A., Boursier, E. and Flammarion, N. (2026) 'Statistical Benefits of Fine-Tuning from Pretrained Initialization in Diagonal Linear Networks', Available at: https://omanscience.com/en/articles/statistical-benefits-of-fine-tuning-from-pretrained-initialization-in-diagonal-linear-networks.
Vancouver
Declèves A, Boursier E, Flammarion N. Statistical Benefits of Fine-Tuning from Pretrained Initialization in Diagonal Linear Networks. https://omanscience.com/en/articles/statistical-benefits-of-fine-tuning-from-pretrained-initialization-in-diagonal-linear-networks
IEEE
A. Declèves, E. Boursier, and N. Flammarion, "Statistical Benefits of Fine-Tuning from Pretrained Initialization in Diagonal Linear Networks," https://omanscience.com/en/articles/statistical-benefits-of-fine-tuning-from-pretrained-initialization-in-diagonal-linear-networks.