Abstract
A single multimodal large language model (MLLM) struggles to excel simultaneously at detection, localization, description, and reasoning in multimodal industrial anomaly understanding (MM-IAU). We show that in-domain training does not close this gap. On MMAD, a widely adopted MM-IAU benchmark, trained specialists reach at most 75.5% accuracy in defect localization, against 92.3% for human experts, and even detect anomalies less accurately than their untrained base model. Meanwhile, different MLLMs offer complementary strengths but share this weakness in fine-grained perception, so combining them alone cannot remove it. We therefore propose SiGMA, a spatially grounded multi-agent framework that divides labor between heterogeneous MLLM agents and a dedicated visual defect expert. A multimodal searcher supplies industrial knowledge and normal references, the defect expert turns query-reference comparison into calibrated anomaly evidence, and a label-free reliability controller weighs each source by task-wise competence and query-level evidence quality. SiGMA reaches 85.2% average accuracy on MMAD, 4.0% above the strongest trained specialist and Gemini-2.5-Pro and within 1.5% of human experts. Even with three agents of at most 9B parameters, it reaches 84.4%, and new MLLMs join without retraining.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Zhang, X., Du, D., Zhu, H., Dai, J., Liu, Y., Liu, G., Zhang, Z., & Long, Z. (2026). Is In-Domain Training Enough for Fine-Grained Industrial Anomaly Understanding? https://omanscience.com/en/articles/is-in-domain-training-enough-for-fine-grained-industrial-anomaly-understanding
MLA 9
Zhang, Xingwu, et al. "Is In-Domain Training Enough for Fine-Grained Industrial Anomaly Understanding?" https://omanscience.com/en/articles/is-in-domain-training-enough-for-fine-grained-industrial-anomaly-understanding.
Chicago (author–date)
Zhang, Xingwu, Duanyang Du, Huiling Zhu, Jiayue Dai, Yixiao Liu, Guozhi Liu, Zhihan Zhang, and Zijun Long. 2026. "Is In-Domain Training Enough for Fine-Grained Industrial Anomaly Understanding?" https://omanscience.com/en/articles/is-in-domain-training-enough-for-fine-grained-industrial-anomaly-understanding.
Harvard
Zhang, X., Du, D., Zhu, H., Dai, J., Liu, Y., Liu, G., Zhang, Z. and Long, Z. (2026) 'Is In-Domain Training Enough for Fine-Grained Industrial Anomaly Understanding?', Available at: https://omanscience.com/en/articles/is-in-domain-training-enough-for-fine-grained-industrial-anomaly-understanding.
Vancouver
Zhang X, Du D, Zhu H, Dai J, Liu Y, Liu G, et al. Is In-Domain Training Enough for Fine-Grained Industrial Anomaly Understanding? https://omanscience.com/en/articles/is-in-domain-training-enough-for-fine-grained-industrial-anomaly-understanding
IEEE
X. Zhang, D. Du, H. Zhu, J. Dai, Y. Liu, G. Liu, Z. Zhang, and Z. Long, "Is In-Domain Training Enough for Fine-Grained Industrial Anomaly Understanding?," https://omanscience.com/en/articles/is-in-domain-training-enough-for-fine-grained-industrial-anomaly-understanding.