Abstract
Conversational MLLM agents are increasingly expected to assist in professional workflows, from AI research and engineering design to product management and business operations. Yet this capability remains underexplored: existing benchmarks largely focus on informal, everyday interactions and personal-life scenarios featuring photographic natural images, isolated static artifacts, and recall-oriented questions. In contrast, professional scenarios often involve structured, information-heavy artifacts that undergo frequent revisions and authority updates, and compositional queries requiring reconciliation of many artifact versions while tracking state precisely. To address these challenges, we introduce DSV-Mem, a benchmark for evaluating Dense Stateful Visual Memory. DSV-Mem comprises expert-reviewed scenarios and 1,000 questions across five user-oriented categories (Current State, Past State, Derived State, Change History, and Conflict/Refusal). A Hartley-inspired criterion favors questions with broader visual-evidence inspection demands. We also introduce a generation harness that produces evaluation suites by decoupling state-transition synthesis from conversation filling. Evaluation over 27 configurations spanning frontier and open-weight models and memory management methods reveals that the strongest baseline scores below 45% on DSV-Mem. Analysis surfaces findings: 1) multimodality and information density both contribute to difficulty, but state evolution, particularly the number of governing updates, is the dominant tested factor. Raw conversation/haystack length, OCR, and arithmetic are not the primary bottlenecks; 2) models often fail to verify user premises against prior state updates before answering; 3) increased reasoning effort and memory management methods yield limited gains, whereas state-aware designs prove more effective. The benchmark and code will be publicly released.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Zhong, J., Chaudhry, R., Chen, X., Zhao, T., Xu, L., Xing, Y., & Sankaran, N. (2026). DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents. https://omanscience.com/en/articles/dsv-mem-evaluating-multimodal-memory-in-professional-workflows-for-mllm-agents
MLA 9
Zhong, Jike, et al. "DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents." https://omanscience.com/en/articles/dsv-mem-evaluating-multimodal-memory-in-professional-workflows-for-mllm-agents.
Chicago (author–date)
Zhong, Jike, Ritwick Chaudhry, Xuanbai Chen, Tianchen Zhao, Linghan Xu, Yifan Xing, and Nishant Sankaran. 2026. "DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents." https://omanscience.com/en/articles/dsv-mem-evaluating-multimodal-memory-in-professional-workflows-for-mllm-agents.
Harvard
Zhong, J., Chaudhry, R., Chen, X., Zhao, T., Xu, L., Xing, Y. and Sankaran, N. (2026) 'DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents', Available at: https://omanscience.com/en/articles/dsv-mem-evaluating-multimodal-memory-in-professional-workflows-for-mllm-agents.
Vancouver
Zhong J, Chaudhry R, Chen X, Zhao T, Xu L, Xing Y, et al. DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents. https://omanscience.com/en/articles/dsv-mem-evaluating-multimodal-memory-in-professional-workflows-for-mllm-agents
IEEE
J. Zhong, R. Chaudhry, X. Chen, T. Zhao, L. Xu, Y. Xing, and N. Sankaran, "DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents," https://omanscience.com/en/articles/dsv-mem-evaluating-multimodal-memory-in-professional-workflows-for-mllm-agents.