Abstract
Molecular optimization is inherently iterative: a candidate is proposed, evaluated against several objectives, and revised while preserving a relationship to the source molecule. Most instruction-following models instead emit one edited molecule, forcing validity, property improvement, and similarity control into a single response. We introduce MARCO, an evaluator-grounded reinforcement-learning framework that trains molecular editors on bounded proposal--feedback--revision trajectories. MARCO aggregates shaped turn rewards into an undiscounted trajectory return for group-relative policy optimization. We evaluate two consequences of this training: Same-1 tests the trained policy under a one-response budget, while Same-5 tests whether the same policy can use verifier feedback when up to five responses are available. Across the three-objective MuMOInstruct benchmark, three Qwen backbones, and seen/unseen instruction splits, SFT-initialized MARCO obtains the highest product of property success rate and similarity in every reported primary setting. Same-5 further improves the observed score under the tested budget, while four-objective and public-checkpoint experiments test transfer across constraint sets and initialization regimes.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Fang, S., Wang, Y., Yang, Z., Xu, X., Lu, J., Tan, C., Zhu, T., Zheng, Y., & Qiu, X. (2026). MARCO: Multi-Round Agentic Reinforcement for Conditional Molecular Optimization. https://omanscience.com/en/articles/marco-multi-round-agentic-reinforcement-for-conditional-molecular-optimization
MLA 9
Fang, Shicheng, et al. "MARCO: Multi-Round Agentic Reinforcement for Conditional Molecular Optimization." https://omanscience.com/en/articles/marco-multi-round-agentic-reinforcement-for-conditional-molecular-optimization.
Chicago (author–date)
Fang, Shicheng, Yuxin Wang, Zhuo Yang, Xiaohu Xu, Jiahao Lu, Chuanyuan Tan, Tong Zhu, Yining Zheng, and Xipeng Qiu. 2026. "MARCO: Multi-Round Agentic Reinforcement for Conditional Molecular Optimization." https://omanscience.com/en/articles/marco-multi-round-agentic-reinforcement-for-conditional-molecular-optimization.
Harvard
Fang, S., Wang, Y., Yang, Z., Xu, X., Lu, J., Tan, C., Zhu, T., Zheng, Y. and Qiu, X. (2026) 'MARCO: Multi-Round Agentic Reinforcement for Conditional Molecular Optimization', Available at: https://omanscience.com/en/articles/marco-multi-round-agentic-reinforcement-for-conditional-molecular-optimization.
Vancouver
Fang S, Wang Y, Yang Z, Xu X, Lu J, Tan C, et al. MARCO: Multi-Round Agentic Reinforcement for Conditional Molecular Optimization. https://omanscience.com/en/articles/marco-multi-round-agentic-reinforcement-for-conditional-molecular-optimization
IEEE
S. Fang, Y. Wang, Z. Yang, X. Xu, J. Lu, C. Tan, T. Zhu, Y. Zheng, and X. Qiu, "MARCO: Multi-Round Agentic Reinforcement for Conditional Molecular Optimization," https://omanscience.com/en/articles/marco-multi-round-agentic-reinforcement-for-conditional-molecular-optimization.