Abstract

Persistent textual memory allows language models to carry information across long interactions, but learning what to remember is fundamentally a credit-assignment problem. A memory rewrite may only become useful many steps later, while much of the observed utility may be inherited from information already stored before the rewrite. We introduce Memory Gain Policy Optimization (MGPO), which isolates the incremental value of each memory rewrite by crediting it for its marginal contribution to current and future downstream utility. This turns delayed memory utility into a direct learning signal for optimizing what information should persist. We study MGPO on document-level information extraction, where structured supervision makes the effects of individual memory updates directly measurable. MGPO improves extraction while reducing average memory length by nearly 80% relative to the initial memory policy before optimization. The learned memory policy also supports reuse and transfer across domains, downstream models without further training. These results show that effective memory learning depends not only on preserving useful information, but on identifying which memory updates create lasting incremental value.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Tang, J., Liu, M., & Sarabi, A. (2026). Learning What to Remember: Long-horizon Counterfactual Memory Optimization. https://omanscience.com/en/articles/learning-what-to-remember-long-horizon-counterfactual-memory-optimization

MLA 9

Tang, Jiaming, et al. "Learning What to Remember: Long-horizon Counterfactual Memory Optimization." https://omanscience.com/en/articles/learning-what-to-remember-long-horizon-counterfactual-memory-optimization.

Chicago (author–date)

Tang, Jiaming, Mingyan Liu, and Armin Sarabi. 2026. "Learning What to Remember: Long-horizon Counterfactual Memory Optimization." https://omanscience.com/en/articles/learning-what-to-remember-long-horizon-counterfactual-memory-optimization.

Harvard

Tang, J., Liu, M. and Sarabi, A. (2026) 'Learning What to Remember: Long-horizon Counterfactual Memory Optimization', Available at: https://omanscience.com/en/articles/learning-what-to-remember-long-horizon-counterfactual-memory-optimization.

Vancouver

Tang J, Liu M, Sarabi A. Learning What to Remember: Long-horizon Counterfactual Memory Optimization. https://omanscience.com/en/articles/learning-what-to-remember-long-horizon-counterfactual-memory-optimization

IEEE

J. Tang, M. Liu, and A. Sarabi, "Learning What to Remember: Long-horizon Counterfactual Memory Optimization," https://omanscience.com/en/articles/learning-what-to-remember-long-horizon-counterfactual-memory-optimization.