Abstract
Composed image retrieval (CIR) aims to retrieve a desired target image from a query consisting of a reference image and a modification text. This task exhibits an unusual representational asymmetry: the modification text specifies a transition from the reference state, whereas retrieval candidates depict completed target states. This creates a representation mismatch for zero-shot methods that query pretrained vision-language spaces directly with transformation-oriented language. We study this mismatch and reformulate zero-shot composed image retrieval as target-state reconstruction followed by retrieval. We instantiate this formulation with ASAP-CIR, a training-free framework that reconstructs a static target representation using a frozen multimodal large language model (MLLM). The representation combines multiple holistic descriptions with a variable set of importance-weighted atomic semantics, thereby preserving both overall target identity and fine-grained visual constraints. Retrieval then integrates holistic state alignment, atomic constraint grounding, and calibrated target-state evidence aggregation. A controlled text-only diagnostic shows that target-side static query formulations achieve more reliable retrieval than dynamic composed query formulations, particularly when source-state semantics must be suppressed or transformed. Experiments on FashionIQ, CIRR, and CIRCO further characterize the effectiveness and limitations of this representation principle, with the clearest gains on the multi-target CIRCO benchmark. These results show that how composed intent is represented before retrieval is a consequential design choice, distinct from the choice of retrieval backbone itself.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Zhao, Y., & Feng, S. (2026). From Transformation to Target State: Rethinking Query Representation for Zero-Shot Composed Image Retrieval. https://omanscience.com/en/articles/from-transformation-to-target-state-rethinking-query-representation-for-zero-shot-composed-image-retrieval
MLA 9
Zhao, Yihe, and Songhe Feng. "From Transformation to Target State: Rethinking Query Representation for Zero-Shot Composed Image Retrieval." https://omanscience.com/en/articles/from-transformation-to-target-state-rethinking-query-representation-for-zero-shot-composed-image-retrieval.
Chicago (author–date)
Zhao, Yihe, and Songhe Feng. 2026. "From Transformation to Target State: Rethinking Query Representation for Zero-Shot Composed Image Retrieval." https://omanscience.com/en/articles/from-transformation-to-target-state-rethinking-query-representation-for-zero-shot-composed-image-retrieval.
Harvard
Zhao, Y. and Feng, S. (2026) 'From Transformation to Target State: Rethinking Query Representation for Zero-Shot Composed Image Retrieval', Available at: https://omanscience.com/en/articles/from-transformation-to-target-state-rethinking-query-representation-for-zero-shot-composed-image-retrieval.
Vancouver
Zhao Y, Feng S. From Transformation to Target State: Rethinking Query Representation for Zero-Shot Composed Image Retrieval. https://omanscience.com/en/articles/from-transformation-to-target-state-rethinking-query-representation-for-zero-shot-composed-image-retrieval
IEEE
Y. Zhao, and S. Feng, "From Transformation to Target State: Rethinking Query Representation for Zero-Shot Composed Image Retrieval," https://omanscience.com/en/articles/from-transformation-to-target-state-rethinking-query-representation-for-zero-shot-composed-image-retrieval.