نسخة أولية وصول مفتوح
SpikeCredit: Temporal Credit Carrier for Reinforcement Learning with Sparse Rewards
Reinforcement learning (RL) with sparse rewards is challenging because delayed outcomes provide little guidance about which intermediate computations caused success or failure. We argue that reliable credit assignment requires policy dynamics that preserve and expose credit-relevant information over time, a role we for …