نسخة أولية وصول مفتوح
Recurrent Self-Improvement: Dynamic Cross-Loop On-Policy Distillation for Looped Language Models
Looped Language Models (LoopLMs) offer a parameter efficient approach to scaling reasoning by reusing shared parameters across recurrent computation steps. Despite their promise, effective post-training of LoopLMs remains challenging. Existing approaches either provide reward based supervision that is sparse or costly …