الملخص
Reinforcement learning (RL) with sparse rewards is challenging because delayed outcomes provide little guidance about which intermediate computations caused success or failure. We argue that reliable credit assignment requires policy dynamics that preserve and expose credit-relevant information over time, a role we formalize as Temporal Credit Carriers (TCCs) and that spiking neural networks (SNNs) naturally fulfill through graded membrane traces and event-driven spikes. Based on this hypothesis, we propose SpikeCredit, an SNN-based framework for RL with sparse rewards that first performs task-adaptive TCC selection and then closes the loop between a fast TCC-reading pathway, where self-motion feedback constraint uses local behavior-grounded cues to constrain transition-level credit recovery, and a slow TCC-writing pathway, where credit-targeted trace alignment feeds recovered credit back into the actor to make future TCC dynamics more credit-readable. Across sparse-reward MuJoCo tasks, SpikeCredit improves Last10 return over sparse SNN baselines by +1169% on Ant, +953% on Hopper, +723% on Swimmer, and +1781% on Walker2d, and exceeds the dense-reward baseline on Swimmer by +113%. Mechanistic analyses further show substantially stronger alignment with dense rewards than the sparse SNN baseline. These results position spiking dynamics as credit-preserving substrates for sparse-reward RL.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Yu, Y., Sun, P., Pan, W., Chen, W., Hong, Y., Hao, K., & Jin, Y. (2026). SpikeCredit: Temporal Credit Carrier for Reinforcement Learning with Sparse Rewards. https://omanscience.com/ar/articles/spikecredit-temporal-credit-carrier-for-reinforcement-learning-with-sparse-rewards
MLA 9
Yu, Yingchao, et al. "SpikeCredit: Temporal Credit Carrier for Reinforcement Learning with Sparse Rewards." https://omanscience.com/ar/articles/spikecredit-temporal-credit-carrier-for-reinforcement-learning-with-sparse-rewards.
شيكاغو (المؤلف–التاريخ)
Yu, Yingchao, Pengfei Sun, Wenxuan Pan, Wei Chen, Yitian Hong, Kuangrong Hao, and Yaochu Jin. 2026. "SpikeCredit: Temporal Credit Carrier for Reinforcement Learning with Sparse Rewards." https://omanscience.com/ar/articles/spikecredit-temporal-credit-carrier-for-reinforcement-learning-with-sparse-rewards.
هارفارد
Yu, Y., Sun, P., Pan, W., Chen, W., Hong, Y., Hao, K. and Jin, Y. (2026) 'SpikeCredit: Temporal Credit Carrier for Reinforcement Learning with Sparse Rewards', Available at: https://omanscience.com/ar/articles/spikecredit-temporal-credit-carrier-for-reinforcement-learning-with-sparse-rewards.
فانكوفر
Yu Y, Sun P, Pan W, Chen W, Hong Y, Hao K, et al. SpikeCredit: Temporal Credit Carrier for Reinforcement Learning with Sparse Rewards. https://omanscience.com/ar/articles/spikecredit-temporal-credit-carrier-for-reinforcement-learning-with-sparse-rewards
IEEE
Y. Yu, P. Sun, W. Pan, W. Chen, Y. Hong, K. Hao, and Y. Jin, "SpikeCredit: Temporal Credit Carrier for Reinforcement Learning with Sparse Rewards," https://omanscience.com/ar/articles/spikecredit-temporal-credit-carrier-for-reinforcement-learning-with-sparse-rewards.