الباحثون

Chenfan Qu

المنشورات 2

نسخة أولية وصول مفتوح

Can We Model the Artifacts Explicitly? Disentangle Artifacts via Pairwise Edit Relations for Image Manipulation Localization

Xuekang Zhu, Kaiwen Feng, Rui-Feng Wang وآخرون · 2026

Image Manipulation Localization (IML) is commonly formulated as a fully supervised learning task that estimates the optimal manipulation mask $y$ for a given image $x$. In this work, we first reveal the latent nature of artifacts and thus reinterpret IML as a latent-variable problem, $P(y|x)=\int P(y|z)\,P(z|x)\,dz$, w …

نسخة أولية وصول مفتوح

PolyOCR-Venus: Unified OCR Foundation Models for Text-Centric Visual Intelligence

GuangJian Team, Kaili Huang, Yongshuo Zhang وآخرون · 2026

Optical Character Recognition (OCR) is evolving from plain-text transcription toward general visual intelligence, requiring models to recognize, localize, and reason over textual information in complex visual environments. However, existing OCR systems often excel at only some tasks and struggle to balance recognition, …

المؤلفون المشاركون