Preprint Open access
Multi-Fidelity Policy Gradients Stabilize Data-Scarce Reinforcement Learning
Policy gradient methods for on-policy reinforcement learning (RL) can become unstable when expensive, scarce target-domain data yield noisy gradient estimates. We address this challenge by complementing limited high-fidelity (HF) target-domain data with abundant, cheap, but biased low-fidelity (LF) data, e.g., from a s …