Abstract

As agents continuously improve by generating and revising Skills, the process that discovers and refines those Skills becomes a learnable object in its own right. Task-Skills directly act on task execution, whereas Meta-Skills govern how agents discover and improve future Skills; their value therefore emerges through the subsequent search processes they induce. Existing approaches improve Meta-Skills from observed raw Skill-search trajectories and branch outcomes. However, branch performance entangles the effects of the initial discovery state and the Meta-Skill revision that generated the search process, making it difficult to characterize what a particular revision actually changed, and pushing updates toward revisions that benefit from favorable states rather than those that improve the process. We introduce HMED (Hindsight Meta-Experience Distillation), a mechanism for constructing Meta-Experience for self-improving agents. HMED revisits the completed event from which a revision originates and re-executes the incumbent and revised Meta-Skills from the same restored discovery state, so that the changes associated with the revision can be observed under a shared condition. Each comparison is distilled into a Meta-Experience, a structured record that can be reused by future updates, so that even revisions that are not ultimately retained still contribute a learning signal. Across three interactive agent benchmarks and both open-source and closed-source models, HMED consistently improves Skill discovery performance over strong baselines, shifting Meta-Skill learning beyond branch outcomes toward the consequences of changing the improvement process.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Feng, Q., Huang, Z., Zhu, Y., Zhang, X., & Dou, Q. (2026). Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents. https://omanscience.com/en/articles/learning-from-revision-consequences-hindsight-meta-experience-distillation-for-self-improving-agents

MLA 9

Feng, Qianhan, et al. "Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents." https://omanscience.com/en/articles/learning-from-revision-consequences-hindsight-meta-experience-distillation-for-self-improving-agents.

Chicago (author–date)

Feng, Qianhan, Zhongzhen Huang, Yakun Zhu, Xiaofan Zhang, and Qi Dou. 2026. "Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents." https://omanscience.com/en/articles/learning-from-revision-consequences-hindsight-meta-experience-distillation-for-self-improving-agents.

Harvard

Feng, Q., Huang, Z., Zhu, Y., Zhang, X. and Dou, Q. (2026) 'Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents', Available at: https://omanscience.com/en/articles/learning-from-revision-consequences-hindsight-meta-experience-distillation-for-self-improving-agents.

Vancouver

Feng Q, Huang Z, Zhu Y, Zhang X, Dou Q. Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents. https://omanscience.com/en/articles/learning-from-revision-consequences-hindsight-meta-experience-distillation-for-self-improving-agents

IEEE

Q. Feng, Z. Huang, Y. Zhu, X. Zhang, and Q. Dou, "Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents," https://omanscience.com/en/articles/learning-from-revision-consequences-hindsight-meta-experience-distillation-for-self-improving-agents.