الباحثون

Somayeh Sojoudi

المنشورات 2

نسخة أولية وصول مفتوح

Recursive Self-Improvement through Multi-Agent Self-Supervision

Hyunin Lee, Jinglue Xu, Jeffrey Seely وآخرون · 2026

Recursive self-improvement (RSI) of a model on non-verifiable tasks, such as open-ended research, faces a supervision bottleneck when its outputs exceed what even human experts can reliably assess, leaving the model itself (optimizee) as the best available optimizer and evaluator. However, a single model instance strug …

نسخة أولية وصول مفتوح

How RL Reshapes LLM Reasoning: Transferability, Coverage, and Scaling Laws

Ziheng Cheng, Yixiao Huang, Hanlin Zhu وآخرون · 2026

Recent studies on reinforcement learning (RL) report seemingly conflicting evidence about large language model (LLM) reasoning. Training on mathematics can improve performance in other domains, yet gains in Pass@1 can coincide with lower Pass@$N$ than the base model. This raises a fundamental question: does RL expand a …

المؤلفون المشاركون