Abstract

Image Manipulation Localization (IML) is commonly formulated as a fully supervised learning task that estimates the optimal manipulation mask $y$ for a given image $x$. In this work, we first reveal the latent nature of artifacts and thus reinterpret IML as a latent-variable problem, $P(y|x)=\int P(y|z)\,P(z|x)\,dz$, where $z$ denotes the artifacts. Following this interpretation, we pinpoint the cause for the current IML models' insufficiency as their implicit artifacts modeling strategy, highlighting the necessity of modeling $z$ in an explicit manner. Without direct labels, feature disentanglement is the most appropriate solution for this explicit modeling. Accordingly, we propose a two-stage learning paradigm with the Pairwise Artifacts Learning (PAL) and Standard Localization (SL) phases to estimate $P(z|x)$ and $P(y|z)$ via edit relations. To support our edit-relation-based learning, we further curate EditGroup-45K, a source-anchored dataset organized into edit groups for pair construction. Extensive experiments show that our PAL paradigm yields consistent improvements across diverse IML architectures, and empirical analyses further verify that PAL does capture artifacts explicitly through feature disentanglement. Code and dataset are available at https://github.com/venus-guangjian/PAL

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhu, X., Feng, K., Wang, R. F., Wang, X., Ma, X., Du, B., Jiang, C., Qu, C., Ye, S., Du, X., Feng, W., Liu, J., & Zhou, J. Z. (2026). Can We Model the Artifacts Explicitly? Disentangle Artifacts via Pairwise Edit Relations for Image Manipulation Localization. https://omanscience.com/en/articles/can-we-model-the-artifacts-explicitly-disentangle-artifacts-via-pairwise-edit-relations-for-image-manipulation-localization

MLA 9

Zhu, Xuekang, et al. "Can We Model the Artifacts Explicitly? Disentangle Artifacts via Pairwise Edit Relations for Image Manipulation Localization." https://omanscience.com/en/articles/can-we-model-the-artifacts-explicitly-disentangle-artifacts-via-pairwise-edit-relations-for-image-manipulation-localization.

Chicago (author–date)

Zhu, Xuekang, Kaiwen Feng, Rui-Feng Wang, Xiwen Wang, Xiaochen Ma, Bo Du, Changjiang Jiang, Chenfan Qu, Songyu Ye, Xia Du, Wentao Feng, Jian Liu, and Ji-Zhe Zhou. 2026. "Can We Model the Artifacts Explicitly? Disentangle Artifacts via Pairwise Edit Relations for Image Manipulation Localization." https://omanscience.com/en/articles/can-we-model-the-artifacts-explicitly-disentangle-artifacts-via-pairwise-edit-relations-for-image-manipulation-localization.

Harvard

Zhu, X., Feng, K., Wang, R. F., Wang, X., Ma, X., Du, B., Jiang, C., Qu, C., Ye, S., Du, X., Feng, W., Liu, J. and Zhou, J. Z. (2026) 'Can We Model the Artifacts Explicitly? Disentangle Artifacts via Pairwise Edit Relations for Image Manipulation Localization', Available at: https://omanscience.com/en/articles/can-we-model-the-artifacts-explicitly-disentangle-artifacts-via-pairwise-edit-relations-for-image-manipulation-localization.

Vancouver

Zhu X, Feng K, Wang RF, Wang X, Ma X, Du B, et al. Can We Model the Artifacts Explicitly? Disentangle Artifacts via Pairwise Edit Relations for Image Manipulation Localization. https://omanscience.com/en/articles/can-we-model-the-artifacts-explicitly-disentangle-artifacts-via-pairwise-edit-relations-for-image-manipulation-localization

IEEE

X. Zhu, K. Feng, R. F. Wang, X. Wang, X. Ma, B. Du, C. Jiang, C. Qu, S. Ye, X. Du, W. Feng, J. Liu, and J. Z. Zhou, "Can We Model the Artifacts Explicitly? Disentangle Artifacts via Pairwise Edit Relations for Image Manipulation Localization," https://omanscience.com/en/articles/can-we-model-the-artifacts-explicitly-disentangle-artifacts-via-pairwise-edit-relations-for-image-manipulation-localization.