Preprint Open access
Pointwise or Pairwise: When Do Pairwise Losses Help Reward Learning, Provably?
Pairwise losses are increasingly used for reward learning even when pointwise rewards are observed, with mixed empirical results. When and why do pairwise losses outperform pointwise losses? We study this question in a grouped offline contextual-bandit setting allowing multiple actions per context, capturing many rewar …