Abstract

Modern models increasingly learn through black-box oracles such as humans, optimization solvers, and external tools that provide feedback without exposing their internal mechanisms. A common remedy is to learn an (action-)value function as a control variate. In this paper, we first observe that even an exact action-value function can be arbitrarily far from variance-optimal. We show that this gap arises because the value function minimizes the noise in each action's own gradient term, while an action can still affect the rest of the gradient estimator through shared parameters. A simple unbiased correction, at no extra oracle cost, can still reduce its variance by an arbitrarily large factor. Motivated by this, we then prove that the residual variance can be decomposed exactly by actions with no cross terms. This decomposition yields a closed-form variance-minimizing correction for neural-network parameters, which can be computed by a simple projection. Empirically, our correction consistently reduces the variance left by the value function and improves learning across all tasks. The source code for all experiments is available at https://github.com/Zihao-Kevin/black_box_opt.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhao, Z., Zhang, S., & Wang, K. (2026). Variance-Optimal Control Variates for Learning with Black-box Feedback. https://omanscience.com/en/articles/variance-optimal-control-variates-for-learning-with-black-box-feedback

MLA 9

Zhao, Zihao, et al. "Variance-Optimal Control Variates for Learning with Black-box Feedback." https://omanscience.com/en/articles/variance-optimal-control-variates-for-learning-with-black-box-feedback.

Chicago (author–date)

Zhao, Zihao, Shuhan Zhang, and Kai Wang. 2026. "Variance-Optimal Control Variates for Learning with Black-box Feedback." https://omanscience.com/en/articles/variance-optimal-control-variates-for-learning-with-black-box-feedback.

Harvard

Zhao, Z., Zhang, S. and Wang, K. (2026) 'Variance-Optimal Control Variates for Learning with Black-box Feedback', Available at: https://omanscience.com/en/articles/variance-optimal-control-variates-for-learning-with-black-box-feedback.

Vancouver

Zhao Z, Zhang S, Wang K. Variance-Optimal Control Variates for Learning with Black-box Feedback. https://omanscience.com/en/articles/variance-optimal-control-variates-for-learning-with-black-box-feedback

IEEE

Z. Zhao, S. Zhang, and K. Wang, "Variance-Optimal Control Variates for Learning with Black-box Feedback," https://omanscience.com/en/articles/variance-optimal-control-variates-for-learning-with-black-box-feedback.