Abstract
Long-context reasoning faces two complementary bottlenecks: retaining evidence across long inputs and sustaining computation across many reasoning steps. Existing approaches largely address them separately, with external memory extending access to distant evidence and latent reasoning compressing multi-step computation. We introduce LatentHarness, which unifies memory access and latent reasoning as sequential latent action selection. At each internal step, the model chooses THINK for further computation, RECALL from a fast-weight memory of input evidence and intermediate reasoning states, or EXIT to emit the next token. We train this policy with counterfactual policy distillation, which branches every action for one step and scores its effect on the emitted token. These gains teach the policy when memory is more useful than further reasoning, while gradients through counterfactual recall teach which intermediate states should be retained in memory for future use. Across six general and long-context reasoning benchmarks, LatentHarness at 1.4B improves on the strongest baselines by 2.8% and 10.0% relative, respectively, and runs 5.9x faster than the strongest long-context baseline.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Wang, X., Wang, S., & Liu, B. (2026). LatentHarness: Learning Latent Actions for Memory and Reasoning via Counterfactual Policy Distillation. https://omanscience.com/en/articles/latentharness-learning-latent-actions-for-memory-and-reasoning-via-counterfactual-policy-distillation
MLA 9
Wang, Xiaoqiang, et al. "LatentHarness: Learning Latent Actions for Memory and Reasoning via Counterfactual Policy Distillation." https://omanscience.com/en/articles/latentharness-learning-latent-actions-for-memory-and-reasoning-via-counterfactual-policy-distillation.
Chicago (author–date)
Wang, Xiaoqiang, Suyuchen Wang, and Bang Liu. 2026. "LatentHarness: Learning Latent Actions for Memory and Reasoning via Counterfactual Policy Distillation." https://omanscience.com/en/articles/latentharness-learning-latent-actions-for-memory-and-reasoning-via-counterfactual-policy-distillation.
Harvard
Wang, X., Wang, S. and Liu, B. (2026) 'LatentHarness: Learning Latent Actions for Memory and Reasoning via Counterfactual Policy Distillation', Available at: https://omanscience.com/en/articles/latentharness-learning-latent-actions-for-memory-and-reasoning-via-counterfactual-policy-distillation.
Vancouver
Wang X, Wang S, Liu B. LatentHarness: Learning Latent Actions for Memory and Reasoning via Counterfactual Policy Distillation. https://omanscience.com/en/articles/latentharness-learning-latent-actions-for-memory-and-reasoning-via-counterfactual-policy-distillation
IEEE
X. Wang, S. Wang, and B. Liu, "LatentHarness: Learning Latent Actions for Memory and Reasoning via Counterfactual Policy Distillation," https://omanscience.com/en/articles/latentharness-learning-latent-actions-for-memory-and-reasoning-via-counterfactual-policy-distillation.