الباحثون

Mingyi Li

المنشورات 3

نسخة أولية وصول مفتوح

Closing the Horizon Gap in Policy Optimization for Adversarial MDPs

We consider policy optimization for online episodic tabular Markov decision processes (MDPs) with adversarial losses and bandit feedback. Policy optimization updates the policy locally at each state and avoids optimization over the occupancy-measure polytope, but its existing regret bounds are larger by a factor of the …

المؤلفون المشاركون