Abstract

Long-context reasoning faces two complementary bottlenecks: retaining evidence across long inputs and sustaining computation across many reasoning steps. Existing approaches largely address them separately, with external memory extending access to distant evidence and latent reasoning compressing multi-step computation. We introduce LatentHarness, which unifies memory access and latent reasoning as sequential latent action selection. At each internal step, the model chooses THINK for further computation, RECALL from a fast-weight memory of input evidence and intermediate reasoning states, or EXIT to emit the next token. We train this policy with counterfactual policy distillation, which branches every action for one step and scores its effect on the emitted token. These gains teach the policy when memory is more useful than further reasoning, while gradients through counterfactual recall teach which intermediate states should be retained in memory for future use. Across six general and long-context reasoning benchmarks, LatentHarness at 1.4B improves on the strongest baselines by 2.8% and 10.0% relative, respectively, and runs 5.9x faster than the strongest long-context baseline.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Wang, X., Wang, S., & Liu, B. (2026). LatentHarness: Learning Latent Actions for Memory and Reasoning via Counterfactual Policy Distillation. https://omanscience.com/en/articles/latentharness-learning-latent-actions-for-memory-and-reasoning-via-counterfactual-policy-distillation

MLA 9

Wang, Xiaoqiang, et al. "LatentHarness: Learning Latent Actions for Memory and Reasoning via Counterfactual Policy Distillation." https://omanscience.com/en/articles/latentharness-learning-latent-actions-for-memory-and-reasoning-via-counterfactual-policy-distillation.

Chicago (author–date)

Wang, Xiaoqiang, Suyuchen Wang, and Bang Liu. 2026. "LatentHarness: Learning Latent Actions for Memory and Reasoning via Counterfactual Policy Distillation." https://omanscience.com/en/articles/latentharness-learning-latent-actions-for-memory-and-reasoning-via-counterfactual-policy-distillation.

Harvard

Wang, X., Wang, S. and Liu, B. (2026) 'LatentHarness: Learning Latent Actions for Memory and Reasoning via Counterfactual Policy Distillation', Available at: https://omanscience.com/en/articles/latentharness-learning-latent-actions-for-memory-and-reasoning-via-counterfactual-policy-distillation.

Vancouver

Wang X, Wang S, Liu B. LatentHarness: Learning Latent Actions for Memory and Reasoning via Counterfactual Policy Distillation. https://omanscience.com/en/articles/latentharness-learning-latent-actions-for-memory-and-reasoning-via-counterfactual-policy-distillation

IEEE

X. Wang, S. Wang, and B. Liu, "LatentHarness: Learning Latent Actions for Memory and Reasoning via Counterfactual Policy Distillation," https://omanscience.com/en/articles/latentharness-learning-latent-actions-for-memory-and-reasoning-via-counterfactual-policy-distillation.