Abstract

Off-policy evaluation (OPE) for contextual bandit policies becomes challenging when action-level importance weighting incurs excessive variance. Doubly robust (DR) estimation remains unbiased under common support but retains these high-variance action-level weights. A prior estimator, Off-policy evaluation with Conjunct Effect Model (OffCEM), replaces them with more stable cluster-level weights, at the cost of relying on local correctness of the reward model. In this paper, we show that, under the assumptions required by DR and OffCEM, there exists an unbiased family of estimators that interpolates between OffCEM and DR. Building on this result, we propose the Variance Optimal-CEM (VOCEM) estimator, which selects the interpolation coefficient to minimize variance. We derive the population-optimal coefficient in closed form and show that the resulting estimator has variance no larger than either endpoint, OffCEM or DR. Experiments in controlled synthetic settings and on two large-action benchmarks show that VOCEM improves upon both endpoints in all 23 evaluated conditions, exhibiting greater stability and empirical robustness.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Felicioni, N., Benigni, M., Dacrema, M. F., & Cremonesi, P. (2026). Variance-Optimal Off-Policy Evaluation with Conjunct Effect Modeling. https://omanscience.com/en/articles/variance-optimal-off-policy-evaluation-with-conjunct-effect-modeling

MLA 9

Felicioni, Nicolò, et al. "Variance-Optimal Off-Policy Evaluation with Conjunct Effect Modeling." https://omanscience.com/en/articles/variance-optimal-off-policy-evaluation-with-conjunct-effect-modeling.

Chicago (author–date)

Felicioni, Nicolò, Michael Benigni, Maurizio Ferrari Dacrema, and Paolo Cremonesi. 2026. "Variance-Optimal Off-Policy Evaluation with Conjunct Effect Modeling." https://omanscience.com/en/articles/variance-optimal-off-policy-evaluation-with-conjunct-effect-modeling.

Harvard

Felicioni, N., Benigni, M., Dacrema, M. F. and Cremonesi, P. (2026) 'Variance-Optimal Off-Policy Evaluation with Conjunct Effect Modeling', Available at: https://omanscience.com/en/articles/variance-optimal-off-policy-evaluation-with-conjunct-effect-modeling.

Vancouver

Felicioni N, Benigni M, Dacrema MF, Cremonesi P. Variance-Optimal Off-Policy Evaluation with Conjunct Effect Modeling. https://omanscience.com/en/articles/variance-optimal-off-policy-evaluation-with-conjunct-effect-modeling

IEEE

N. Felicioni, M. Benigni, M. F. Dacrema, and P. Cremonesi, "Variance-Optimal Off-Policy Evaluation with Conjunct Effect Modeling," https://omanscience.com/en/articles/variance-optimal-off-policy-evaluation-with-conjunct-effect-modeling.