الملخص

Despite significant progress in visual tasks by Multimodal Large Language Models (MLLMs), geometric diagram understanding remains challenging due to the presence of sparse visual cues and ambiguous symbol-primitive associations. MLLMs may therefore rely on textual priors, producing interpretations that conflict with visual evidence. We introduce the training-free Criticality-Driven Visual Intervention Framework (CVIF), an inference-time method that localizes critical layers and executes visual interventions during the transition from evidence aggregation to semantic decoding. At these layers, a Geometry-Constrained Local Relation Reconstruction (GCLR) module selects and weights vertex-centered visual evidence, while an Adaptive Visual Steering Operator (AVSO) redistributes attention mass toward the selected tokens. Experiments on PGPS9K and PGDP5K show that CVIF raises Overall F1 from 77.85 to 85.58 and from 75.23 to 82.84, respectively, establishing a novel inference-time visual intervention paradigm.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Kang, J., Wei, B., Zhang, L., Jiang, T., Xiao, Q., Zhang, J., & Liu, J. (2026). CVIF: A Criticality-Driven Visual Intervention Framework for Geometric Diagram Understanding in MLLMs. https://omanscience.com/ar/articles/cvif-a-criticality-driven-visual-intervention-framework-for-geometric-diagram-understanding-in-mllms

MLA 9

Kang, Jiahui, et al. "CVIF: A Criticality-Driven Visual Intervention Framework for Geometric Diagram Understanding in MLLMs." https://omanscience.com/ar/articles/cvif-a-criticality-driven-visual-intervention-framework-for-geometric-diagram-understanding-in-mllms.

شيكاغو (المؤلف–التاريخ)

Kang, Jiahui, Bifan Wei, Lingling Zhang, Tianwen Jiang, Qiuyong Xiao, Jihong Zhang, and Jun Liu. 2026. "CVIF: A Criticality-Driven Visual Intervention Framework for Geometric Diagram Understanding in MLLMs." https://omanscience.com/ar/articles/cvif-a-criticality-driven-visual-intervention-framework-for-geometric-diagram-understanding-in-mllms.

هارفارد

Kang, J., Wei, B., Zhang, L., Jiang, T., Xiao, Q., Zhang, J. and Liu, J. (2026) 'CVIF: A Criticality-Driven Visual Intervention Framework for Geometric Diagram Understanding in MLLMs', Available at: https://omanscience.com/ar/articles/cvif-a-criticality-driven-visual-intervention-framework-for-geometric-diagram-understanding-in-mllms.

فانكوفر

Kang J, Wei B, Zhang L, Jiang T, Xiao Q, Zhang J, et al. CVIF: A Criticality-Driven Visual Intervention Framework for Geometric Diagram Understanding in MLLMs. https://omanscience.com/ar/articles/cvif-a-criticality-driven-visual-intervention-framework-for-geometric-diagram-understanding-in-mllms

IEEE

J. Kang, B. Wei, L. Zhang, T. Jiang, Q. Xiao, J. Zhang, and J. Liu, "CVIF: A Criticality-Driven Visual Intervention Framework for Geometric Diagram Understanding in MLLMs," https://omanscience.com/ar/articles/cvif-a-criticality-driven-visual-intervention-framework-for-geometric-diagram-understanding-in-mllms.