Abstract
Reinforcement learning for large language models typically maximizes expected return, adding up the probabilities of all successful trajectories. However, the classical sum formulation can only report how often the model policy succeeds, not which solution actually worked, and because probabilities sum to one, reinforcing one solution can make the model forget another that was never shown to be wrong. This makes expected return a poor fit for compositional reasoning, where a solution must be assembled from reasoning steps that the model produces in separate, often failed, attempts but rarely produces together. To address this, we propose Tropical Reinforcement Learning, which rests on a simple change of algebra: instead of adding the probabilities of alternative solutions, we take their maximum, which yields the tropical semiring. The value of a state then becomes the log-probability of its most likely verified solution, together with an explicit path that can be replayed and reused. This enables true composition, since the best prefix and the best suffix meeting at a shared state can be joined even when they come from different rollouts. To put this into practice, we introduce TROPIC, a training algorithm for deterministic, resettable environments with verifiable outcomes. On four agentic tasks (Sokoban, Countdown, FrozenLake, WebShop), TROPIC outperforms the strongest on-policy baselines by up to 16 percentage points. Changing the algebra of reinforcement learning, not just its estimators, can thus substantially improve compositional reasoning in language models
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Asadulaev, A., Djuhera, A., Salta, K., Boche, H., Karray, F., & Takac, M. (2026). Tropical Reinforcement Learning. https://omanscience.com/en/articles/tropical-reinforcement-learning
MLA 9
Asadulaev, Arip, et al. "Tropical Reinforcement Learning." https://omanscience.com/en/articles/tropical-reinforcement-learning.
Chicago (author–date)
Asadulaev, Arip, Aladin Djuhera, Karim Salta, Holger Boche, Fakhri Karray, and Martin Takac. 2026. "Tropical Reinforcement Learning." https://omanscience.com/en/articles/tropical-reinforcement-learning.
Harvard
Asadulaev, A., Djuhera, A., Salta, K., Boche, H., Karray, F. and Takac, M. (2026) 'Tropical Reinforcement Learning', Available at: https://omanscience.com/en/articles/tropical-reinforcement-learning.
Vancouver
Asadulaev A, Djuhera A, Salta K, Boche H, Karray F, Takac M. Tropical Reinforcement Learning. https://omanscience.com/en/articles/tropical-reinforcement-learning
IEEE
A. Asadulaev, A. Djuhera, K. Salta, H. Boche, F. Karray, and M. Takac, "Tropical Reinforcement Learning," https://omanscience.com/en/articles/tropical-reinforcement-learning.