[
    {
        "id": "osp-15590",
        "type": "article-journal",
        "title": "DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents",
        "author": [
            {
                "family": "Zhong",
                "given": "Jike"
            },
            {
                "family": "Chaudhry",
                "given": "Ritwick"
            },
            {
                "family": "Chen",
                "given": "Xuanbai"
            },
            {
                "family": "Zhao",
                "given": "Tianchen"
            },
            {
                "family": "Xu",
                "given": "Linghan"
            },
            {
                "family": "Xing",
                "given": "Yifan"
            },
            {
                "family": "Sankaran",
                "given": "Nishant"
            }
        ],
        "URL": "https://omanscience.com/en/articles/dsv-mem-evaluating-multimodal-memory-in-professional-workflows-for-mllm-agents",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Conversational MLLM agents are increasingly expected to assist in professional workflows, from AI research and engineering design to product management and business operations. Yet this capability remains underexplored: existing benchmarks largely focus on informal, everyday interactions and personal-life scenarios featuring photographic natural images, isolated static artifacts, and recall-oriented questions. In contrast, professional scenarios often involve structured, information-heavy artifacts that undergo frequent revisions and authority updates, and compositional queries requiring reconciliation of many artifact versions while tracking state precisely. To address these challenges, we introduce DSV-Mem, a benchmark for evaluating Dense Stateful Visual Memory. DSV-Mem comprises expert-reviewed scenarios and 1,000 questions across five user-oriented categories (Current State, Past State, Derived State, Change History, and Conflict/Refusal). A Hartley-inspired criterion favors questions with broader visual-evidence inspection demands. We also introduce a generation harness that produces evaluation suites by decoupling state-transition synthesis from conversation filling. Evaluation over 27 configurations spanning frontier and open-weight models and memory management methods reveals that the strongest baseline scores below 45% on DSV-Mem. Analysis surfaces findings: 1) multimodality and information density both contribute to difficulty, but state evolution, particularly the number of governing updates, is the dominant tested factor. Raw conversation/haystack length, OCR, and arithmetic are not the primary bottlenecks; 2) models often fail to verify user premises against prior state updates before answering; 3) increased reasoning effort and memory management methods yield limited gains, whereas state-aware designs prove more effective. The benchmark and code will be publicly released."
    }
]