Abstract

Offline multi-agent reinforcement learning (MARL) faces a persistent trade-off. Expressive generative policies can represent multi-modal coordination in the data, but cannot distinguish high-value regions, while value-optimized policies exploit the learned Q-function but collapse the multi-modal into a single dominant mode. A single agent's mode collapse can break joint coordination, and simultaneous drift across agents can push the joint policy into unseen regions of the action space. We propose scalable coordination via optimal unified transport (SCOUT), the first offline MARL framework to combine a generative foundation model with a learned value function through test-time action refinement. SCOUT trains two decoupled components: a flow-matching behavioral prior and a decomposed value function. At test-time, it transports behavioral samples toward high-value regions via Stein variational gradient descent. The number of transport steps controls adaptive test-time scaling, replacing a fixed regularization coefficient. Under the individual-global-max (IGM) principle, we prove a single-term KL bound on the joint soft-value gap that vanishes as transport converges, with an irreducible additive residual proportional to the IGM violation. Empirically, SCOUT achieves the best average performance across discrete and continuous offline MARL benchmarks and yields performance improvements in all offline-to-online configurations.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Lee, D., Xu, H., & Zhang, A. (2026). Test-time Multi-agent Coordination by Decomposed Value Gradient Flow. https://omanscience.com/en/articles/test-time-multi-agent-coordination-by-decomposed-value-gradient-flow

MLA 9

Lee, Dongsu, et al. "Test-time Multi-agent Coordination by Decomposed Value Gradient Flow." https://omanscience.com/en/articles/test-time-multi-agent-coordination-by-decomposed-value-gradient-flow.

Chicago (author–date)

Lee, Dongsu, Haoran Xu, and Amy Zhang. 2026. "Test-time Multi-agent Coordination by Decomposed Value Gradient Flow." https://omanscience.com/en/articles/test-time-multi-agent-coordination-by-decomposed-value-gradient-flow.

Harvard

Lee, D., Xu, H. and Zhang, A. (2026) 'Test-time Multi-agent Coordination by Decomposed Value Gradient Flow', Available at: https://omanscience.com/en/articles/test-time-multi-agent-coordination-by-decomposed-value-gradient-flow.

Vancouver

Lee D, Xu H, Zhang A. Test-time Multi-agent Coordination by Decomposed Value Gradient Flow. https://omanscience.com/en/articles/test-time-multi-agent-coordination-by-decomposed-value-gradient-flow

IEEE

D. Lee, H. Xu, and A. Zhang, "Test-time Multi-agent Coordination by Decomposed Value Gradient Flow," https://omanscience.com/en/articles/test-time-multi-agent-coordination-by-decomposed-value-gradient-flow.