Abstract
Large language models (LLMs) exhibit strong general capabilities that mechanistic interpretability has attributed to sparse computational circuits. However, existing circuit studies emphasize preserving functionality or explaining safety, leaving the mechanisms underlying failures across a broader range of tasks largely unexplored. Extending circuit analysis from abilities to errors, we explore the perspective that such failures may likewise arise from erroneous internal computations and that targeted tuning of the corresponding parameters can correct such errors while largely preserving other capabilities. Motivated by this insight, we introduce RESCUE (Reasoning-Error Sparse-Circuit Uncovering and Editing), a framework that localizes error-associated circuits and surgically repairs them for performance enhancement. General tasks typically involve multi-step reasoning and long-form generation, where early deviations can cause prefixes to drift from supervised references, leading SFT-based mask optimization to overlook circuits involved in generation-time errors. RESCUE therefore refines these masks through reinforcement learning with multiple masked-model rollouts, improving their relevance to observed task failures. Finally, RESCUE introduces a pruning technique and precisely fine-tunes error circuits to correct task failures, thereby translating error localization into a sparse and targeted model update. We validate RESCUE on heterogeneous repair sets across two domains: (1) mathematical reasoning, identifying a math error circuit of 1.40% density whose repair raises accuracy from 6.0% to 75.5%; and (2) medical QA, where a similarly compact 1.44% circuit improves repair-set accuracy from 0% to 81%. Our code is available at: https://github.com/chuanpupig/RESCUE.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Liu, C., Yu, M., Cai, Y., Zhang, Y., Zhou, Z., Sun, L., Jiang, Z., & Guo, Y. (2026). RESCUE: Repairing Language Model Errors to Sparse Circuits via Reinforcement Learning. https://omanscience.com/en/articles/rescue-repairing-language-model-errors-to-sparse-circuits-via-reinforcement-learning
MLA 9
Liu, Chuanpu, et al. "RESCUE: Repairing Language Model Errors to Sparse Circuits via Reinforcement Learning." https://omanscience.com/en/articles/rescue-repairing-language-model-errors-to-sparse-circuits-via-reinforcement-learning.
Chicago (author–date)
Liu, Chuanpu, Miao Yu, Yikai Cai, Yuanhe Zhang, Zhenhong Zhou, Li Sun, Zuming Jiang, and Yufei Guo. 2026. "RESCUE: Repairing Language Model Errors to Sparse Circuits via Reinforcement Learning." https://omanscience.com/en/articles/rescue-repairing-language-model-errors-to-sparse-circuits-via-reinforcement-learning.
Harvard
Liu, C., Yu, M., Cai, Y., Zhang, Y., Zhou, Z., Sun, L., Jiang, Z. and Guo, Y. (2026) 'RESCUE: Repairing Language Model Errors to Sparse Circuits via Reinforcement Learning', Available at: https://omanscience.com/en/articles/rescue-repairing-language-model-errors-to-sparse-circuits-via-reinforcement-learning.
Vancouver
Liu C, Yu M, Cai Y, Zhang Y, Zhou Z, Sun L, et al. RESCUE: Repairing Language Model Errors to Sparse Circuits via Reinforcement Learning. https://omanscience.com/en/articles/rescue-repairing-language-model-errors-to-sparse-circuits-via-reinforcement-learning
IEEE
C. Liu, M. Yu, Y. Cai, Y. Zhang, Z. Zhou, L. Sun, Z. Jiang, and Y. Guo, "RESCUE: Repairing Language Model Errors to Sparse Circuits via Reinforcement Learning," https://omanscience.com/en/articles/rescue-repairing-language-model-errors-to-sparse-circuits-via-reinforcement-learning.