الباحثون

Holger Boche

المنشورات 1

نسخة أولية وصول مفتوح

Tropical Reinforcement Learning

Arip Asadulaev, Aladin Djuhera, Karim Salta وآخرون · 2026

Reinforcement learning for large language models typically maximizes expected return, adding up the probabilities of all successful trajectories. However, the classical sum formulation can only report how often the model policy succeeds, not which solution actually worked, and because probabilities sum to one, reinforc …

المؤلفون المشاركون