Abstract

Reinforcement learning has greatly advanced the capabilities of large language models, but its memory demands remain a barrier to broader adoption. We introduce LoGRA, an approach to RL post-training that reduces memory by retaining useful learning signals in low-rank gradient sketches. These compact representations support both model updates and efficient policy synchronization. To prevent overly large updates from disrupting learning, we complement gradient compression with predicted-KL step control, which estimates policy changes before applying each update and adjusts its magnitude accordingly. With all techniques combined, LoGRA reduces average training memory usage by up to 45.7% across reasoning tasks without compromising performance. It also enables stable training of a 27B-parameter model for over 1,100 steps on a single eight-GPU node, where dense Adam runs out of memory, making previously memory-infeasible RL training practical. Code is available in the \href{https://github.com/skzhang1/labs-molt/tree/logra/examples/scripts/logra}{Molt library}.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhang, S., Zhang, Y., Hu, J., Li, Y., Zhang, H., Xu, B., Kautz, J., & Dong, Y. (2026). LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches. https://omanscience.com/en/articles/logra-scaling-llm-reinforcement-learning-with-low-rank-gradient-sketches

MLA 9

Zhang, Shaokun, et al. "LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches." https://omanscience.com/en/articles/logra-scaling-llm-reinforcement-learning-with-low-rank-gradient-sketches.

Chicago (author–date)

Zhang, Shaokun, Yifan Zhang, Jian Hu, Yueying Li, Hao Zhang, Binfeng Xu, Jan Kautz, and Yi Dong. 2026. "LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches." https://omanscience.com/en/articles/logra-scaling-llm-reinforcement-learning-with-low-rank-gradient-sketches.

Harvard

Zhang, S., Zhang, Y., Hu, J., Li, Y., Zhang, H., Xu, B., Kautz, J. and Dong, Y. (2026) 'LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches', Available at: https://omanscience.com/en/articles/logra-scaling-llm-reinforcement-learning-with-low-rank-gradient-sketches.

Vancouver

Zhang S, Zhang Y, Hu J, Li Y, Zhang H, Xu B, et al. LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches. https://omanscience.com/en/articles/logra-scaling-llm-reinforcement-learning-with-low-rank-gradient-sketches

IEEE

S. Zhang, Y. Zhang, J. Hu, Y. Li, H. Zhang, B. Xu, J. Kautz, and Y. Dong, "LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches," https://omanscience.com/en/articles/logra-scaling-llm-reinforcement-learning-with-low-rank-gradient-sketches.