Abstract

Memory has become integral to the LLM agent ecosystem, supporting information retention and reuse across interactions. However, most existing agent memory systems construct memory in a query-agnostic manner, which can incur unnecessary preprocessing cost and discard details that later prove essential. Recent studies have begun shifting memory processing toward runtime adaptation, but typically specialize in particular operations or fixed processing schemes, leaving flexible control over performance, cost, and latency largely underexplored. To address this challenge, we present \textbf{MemPilot}, a flexible framework that orchestrates on-demand memory curation under different performance--cost--latency preferences. Specifically, we optimize a multi-step LLM policy via reinforcement learning to iteratively choose between retrieving from query-agnostic memory and delegating query-specific curation of raw multimodal history to heterogeneous LLMs and VLMs. The policy jointly controls evidence amount, curation instructions, model selection, and visual access, enabling fine-grained allocation of runtime computation. To optimize this policy under competing objectives, we adapt objective-wise advantage decoupling by separately estimating each objective's advantage before aggregation. Moreover, we introduce prefix-based marginal utility estimation for fine-grained credit assignment across multi-step rollouts. Experiments on five multimodal agent-memory benchmarks demonstrate favorable performance--cost--latency trade-offs across optimization preferences, with preference sweeps yielding broader frontiers than existing trade-off-aware baselines.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhang, H., Yue, H., Long, Q., Bao, J., Liu, Q., Feng, T., Liu, B., Liang, W., & Wang, W. (2026). MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents. https://omanscience.com/en/articles/mempilot-orchestrating-on-demand-multimodal-memory-curation-for-llm-agents

MLA 9

Zhang, Haozhen, et al. "MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents." https://omanscience.com/en/articles/mempilot-orchestrating-on-demand-multimodal-memory-curation-for-llm-agents.

Chicago (author–date)

Zhang, Haozhen, Haodong Yue, Quanyu Long, Jianzhu Bao, Qingyuan Liu, Tao Feng, Bohan Liu, Weida Liang, and Wenya Wang. 2026. "MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents." https://omanscience.com/en/articles/mempilot-orchestrating-on-demand-multimodal-memory-curation-for-llm-agents.

Harvard

Zhang, H., Yue, H., Long, Q., Bao, J., Liu, Q., Feng, T., Liu, B., Liang, W. and Wang, W. (2026) 'MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents', Available at: https://omanscience.com/en/articles/mempilot-orchestrating-on-demand-multimodal-memory-curation-for-llm-agents.

Vancouver

Zhang H, Yue H, Long Q, Bao J, Liu Q, Feng T, et al. MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents. https://omanscience.com/en/articles/mempilot-orchestrating-on-demand-multimodal-memory-curation-for-llm-agents

IEEE

H. Zhang, H. Yue, Q. Long, J. Bao, Q. Liu, T. Feng, B. Liu, W. Liang, and W. Wang, "MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents," https://omanscience.com/en/articles/mempilot-orchestrating-on-demand-multimodal-memory-curation-for-llm-agents.