Abstract

Multi-agent reinforcement learning (MARL) provides a powerful framework for learning coordinated behaviors through interactions with the environment. Developing MARL policies requires balancing expressive modeling of complex and multimodal action distributions with efficient training and execution. Generative policies, particularly diffusionbased policies, can faithfully capture complex and multimodal behaviors, but costly iterative sampling hinders their scalability in online multi-agent settings. We propose an Online MARL framework via one-step Flow model (OMAF) that combines expressive generative policies with efficient one-step action generation. OMAF employs a Transformer-based flow policy to capture complex coordination behaviors, while its approximate path score surrogate provides a principled route to synchronized flow policy optimization. To enable stable and sampleefficient learning, we further develop a joint optimization scheme coupling softmax Q-value estimation with a joint flow policy objective for coordinated policy learning. By eliminating iterative sampling, OMAF dramatically reduces training overhead without sacrificing policy expressiveness. Extensive experiments across 10 standard tasks from MPE and MAMuJoCo show that OMAF consistently achieves superior performance, with up to 3.4x higher returns and 10.5x sample efficiency improvement compared with baseline methods. These results validate the effectiveness of OMAF as an expressive and computationally efficient one-step flow policy paradigm for online MARL.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Li, Z., Li, Y., Wang, X., Du, Y., & Huang, L. (2026). Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies. https://omanscience.com/en/articles/flowing-faster-to-coordinate-one-step-online-multi-agent-flow-policies

MLA 9

Li, Zhuoran, et al. "Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies." https://omanscience.com/en/articles/flowing-faster-to-coordinate-one-step-online-multi-agent-flow-policies.

Chicago (author–date)

Li, Zhuoran, Yunzhan Li, Xun Wang, Yihan Du, and Longbo Huang. 2026. "Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies." https://omanscience.com/en/articles/flowing-faster-to-coordinate-one-step-online-multi-agent-flow-policies.

Harvard

Li, Z., Li, Y., Wang, X., Du, Y. and Huang, L. (2026) 'Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies', Available at: https://omanscience.com/en/articles/flowing-faster-to-coordinate-one-step-online-multi-agent-flow-policies.

Vancouver

Li Z, Li Y, Wang X, Du Y, Huang L. Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies. https://omanscience.com/en/articles/flowing-faster-to-coordinate-one-step-online-multi-agent-flow-policies

IEEE

Z. Li, Y. Li, X. Wang, Y. Du, and L. Huang, "Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies," https://omanscience.com/en/articles/flowing-faster-to-coordinate-one-step-online-multi-agent-flow-policies.