الباحثون

Junran Wang

المنشورات 3

نسخة أولية وصول مفتوح

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exa …

المؤلفون المشاركون