الملخص
Pre-pretraining (PPT) on synthetic non-natural language data improves token efficiency during language model pre-training (PT). Prior work attributes this gain to a grammatical prior, i.e., a structural inductive bias learned during PPT that transfers to natural language grammar. However, PPT has only been tested on models of at most 1B parameters and PT budgets below 2B tokens on predominantly web text. It is unknown whether PPT is effective at larger scales and under more realistic PT data mixtures that combine diverse sources (e.g., code and math). We therefore present a comprehensive study on PPT spanning five PPT tasks, four PT data mixtures, four parameter scales (500M to 7B), and PT budgets of up to 100B tokens. Our results demonstrate that the downstream performance and token efficiency gains of PPT persist at scale, e.g., saving at least 21B PT tokens at the 3B scale. However, in contrast to prior work, we find no consistent evidence that these gains stem from a grammatical prior. Downstream performance does not consistently align with grammatical acceptability across model sizes. Instead, we find that downstream gains arise from PPT tasks that improve long-range retrieval. Finally, PPT performance gains are robust to how PT data mixtures are composed and diminish only when web text is absent. Overall, PPT is a low-cost addition to PT, and future PPT task design should target long-range retrieval rather than natural language grammar.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Yamaguchi, A., Inaba, T., Niklaus, J., Štefánik, M., Villavicencio, A., & Aletras, N. (2026). Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior. https://omanscience.com/ar/articles/synthetic-pre-pretraining-survives-scale-but-not-as-a-grammatical-prior
MLA 9
Yamaguchi, Atsuki, et al. "Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior." https://omanscience.com/ar/articles/synthetic-pre-pretraining-survives-scale-but-not-as-a-grammatical-prior.
شيكاغو (المؤلف–التاريخ)
Yamaguchi, Atsuki, Tatsuro Inaba, Joel Niklaus, Michal Štefánik, Aline Villavicencio, and Nikolaos Aletras. 2026. "Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior." https://omanscience.com/ar/articles/synthetic-pre-pretraining-survives-scale-but-not-as-a-grammatical-prior.
هارفارد
Yamaguchi, A., Inaba, T., Niklaus, J., Štefánik, M., Villavicencio, A. and Aletras, N. (2026) 'Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior', Available at: https://omanscience.com/ar/articles/synthetic-pre-pretraining-survives-scale-but-not-as-a-grammatical-prior.
فانكوفر
Yamaguchi A, Inaba T, Niklaus J, Štefánik M, Villavicencio A, Aletras N. Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior. https://omanscience.com/ar/articles/synthetic-pre-pretraining-survives-scale-but-not-as-a-grammatical-prior
IEEE
A. Yamaguchi, T. Inaba, J. Niklaus, M. Štefánik, A. Villavicencio, and N. Aletras, "Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior," https://omanscience.com/ar/articles/synthetic-pre-pretraining-survives-scale-but-not-as-a-grammatical-prior.