Abstract

Traditional reinforcement learning (RL) techniques focus on maximizing expected cumulative reward, where each action assumes to take a constant unit of time. However, this assumption does not hold for agentic RL tasks such as machine learning engineering (MLE) agents, where actions involve data loading, feature engineering, and model training that take variable durations. Efficiency matters in modern agentic RL where actions are costly. To address this limitation, we adapt from continuous-time RL and Semi-Markov Decision Process (SMDP) formulation and propose Reward-rate Policy Gradient (RPG), where we focus on optimizing the reward rate -- the long-term reward per unit of time. RPG estimates the reward rate from off-policy samples, then charges each action for the time it consumes at that rate. We first conduct theoretical analysis in the bandit setting to establish that RPG approximates the optimal reward rate and empirically demonstrate it outperforms baselines while avoiding enumeration over the policy space, a known issue for an existing method. We then further apply RPG on a small language model (Qwen3.5-4B) with self-improvement loops and empirically show it obtains higher rewards within a fixed time budget than vanilla RL on MLE-Bench and NanoGPT, with a 19.2% and 85.7% margin, respectively. Our method provides a practical solution for optimizing performance under wait time considerations in modern agentic RL tasks, where actions interact with external environments and cost time.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Tian, M., & Yang, S. (2026). Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents. https://omanscience.com/en/articles/reward-rate-policy-gradient-for-efficient-machine-learning-engineering-agents

MLA 9

Tian, Muhang, and Sherry Yang. "Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents." https://omanscience.com/en/articles/reward-rate-policy-gradient-for-efficient-machine-learning-engineering-agents.

Chicago (author–date)

Tian, Muhang, and Sherry Yang. 2026. "Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents." https://omanscience.com/en/articles/reward-rate-policy-gradient-for-efficient-machine-learning-engineering-agents.

Harvard

Tian, M. and Yang, S. (2026) 'Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents', Available at: https://omanscience.com/en/articles/reward-rate-policy-gradient-for-efficient-machine-learning-engineering-agents.

Vancouver

Tian M, Yang S. Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents. https://omanscience.com/en/articles/reward-rate-policy-gradient-for-efficient-machine-learning-engineering-agents

IEEE

M. Tian, and S. Yang, "Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents," https://omanscience.com/en/articles/reward-rate-policy-gradient-for-efficient-machine-learning-engineering-agents.