الباحثون

Joel Niklaus

المنشورات 1

نسخة أولية وصول مفتوح

Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior

Pre-pretraining (PPT) on synthetic non-natural language data improves token efficiency during language model pre-training (PT). Prior work attributes this gain to a grammatical prior, i.e., a structural inductive bias learned during PPT that transfers to natural language grammar. However, PPT has only been tested on mo …

المؤلفون المشاركون