Abstract

Reinforcement learning for large language models typically maximizes expected return, adding up the probabilities of all successful trajectories. However, the classical sum formulation can only report how often the model policy succeeds, not which solution actually worked, and because probabilities sum to one, reinforcing one solution can make the model forget another that was never shown to be wrong. This makes expected return a poor fit for compositional reasoning, where a solution must be assembled from reasoning steps that the model produces in separate, often failed, attempts but rarely produces together. To address this, we propose Tropical Reinforcement Learning, which rests on a simple change of algebra: instead of adding the probabilities of alternative solutions, we take their maximum, which yields the tropical semiring. The value of a state then becomes the log-probability of its most likely verified solution, together with an explicit path that can be replayed and reused. This enables true composition, since the best prefix and the best suffix meeting at a shared state can be joined even when they come from different rollouts. To put this into practice, we introduce TROPIC, a training algorithm for deterministic, resettable environments with verifiable outcomes. On four agentic tasks (Sokoban, Countdown, FrozenLake, WebShop), TROPIC outperforms the strongest on-policy baselines by up to 16 percentage points. Changing the algebra of reinforcement learning, not just its estimators, can thus substantially improve compositional reasoning in language models

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Asadulaev, A., Djuhera, A., Salta, K., Boche, H., Karray, F., & Takac, M. (2026). Tropical Reinforcement Learning. https://omanscience.com/en/articles/tropical-reinforcement-learning

MLA 9

Asadulaev, Arip, et al. "Tropical Reinforcement Learning." https://omanscience.com/en/articles/tropical-reinforcement-learning.

Chicago (author–date)

Asadulaev, Arip, Aladin Djuhera, Karim Salta, Holger Boche, Fakhri Karray, and Martin Takac. 2026. "Tropical Reinforcement Learning." https://omanscience.com/en/articles/tropical-reinforcement-learning.

Harvard

Asadulaev, A., Djuhera, A., Salta, K., Boche, H., Karray, F. and Takac, M. (2026) 'Tropical Reinforcement Learning', Available at: https://omanscience.com/en/articles/tropical-reinforcement-learning.

Vancouver

Asadulaev A, Djuhera A, Salta K, Boche H, Karray F, Takac M. Tropical Reinforcement Learning. https://omanscience.com/en/articles/tropical-reinforcement-learning

IEEE

A. Asadulaev, A. Djuhera, K. Salta, H. Boche, F. Karray, and M. Takac, "Tropical Reinforcement Learning," https://omanscience.com/en/articles/tropical-reinforcement-learning.