Abstract

Text-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count. Recent architectures already augment CLIP with a T5 encoder precisely because CLIP's contrastive embedding loses compositional structure, yet these failures persist. We argue the binding problem is therefore not one of missing information but of misaligned information: a text encoder preserves compositional structure, but in a representation space shaped by language modelling rather than vision, and the denoising objective does not directly reward aligning the two. We show this correspondence can be supplied as an explicit training signal, that the relevant cross-modal information is concentrated in a low-rank subspace of self-supervised visual features, and that supplying it can be folded into diffusion training as a single auxiliary loss. Our method, ORCA (Orthogonal Residual Compositional Alignment), aligns the latent of a diffusion transformer with a low-rank target derived from a frozen visual encoder, through a predictor whose orthogonal basis is parameterised by a learned residual between T5 and CLIP embeddings, which provides a prompt-dependent signal for selecting the visual readout subspace. We prove that the cross-modal information recoverable at a given rank is bounded by the spectral mass of the visual encoder's covariance in the top components. Across three diffusion-transformer backbones (DiT-B/2, DiT-L/2, U-ViT-L), ORCA improves FID and GenEval over both vanilla and REPA baselines at zero inference-time cost; on DiT-L/2 it reaches FID 16.65 and GenEval 0.291 at 200K steps, exceeding the strongest 400K baseline at half the training cost, with the largest gains concentrated on attribute binding, spatial relations, and multi-object prompts.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Hemmat, A., Vahidi, A., Shidani, A., Sanian, M. V., Asadollahzadeh, H., Parast, A. Y., & Lotfollahi, M. (2026). ORCA: Hunting Compositional Failures in Text-to-Image Diffusion. https://omanscience.com/en/articles/orca-hunting-compositional-failures-in-text-to-image-diffusion

MLA 9

Hemmat, Arshia, et al. "ORCA: Hunting Compositional Failures in Text-to-Image Diffusion." https://omanscience.com/en/articles/orca-hunting-compositional-failures-in-text-to-image-diffusion.

Chicago (author–date)

Hemmat, Arshia, Amirhossein Vahidi, Amitis Shidani, Mohammad Vali Sanian, Hesam Asadollahzadeh, Aryan Yazdan Parast, and Mohammad Lotfollahi. 2026. "ORCA: Hunting Compositional Failures in Text-to-Image Diffusion." https://omanscience.com/en/articles/orca-hunting-compositional-failures-in-text-to-image-diffusion.

Harvard

Hemmat, A., Vahidi, A., Shidani, A., Sanian, M. V., Asadollahzadeh, H., Parast, A. Y. and Lotfollahi, M. (2026) 'ORCA: Hunting Compositional Failures in Text-to-Image Diffusion', Available at: https://omanscience.com/en/articles/orca-hunting-compositional-failures-in-text-to-image-diffusion.

Vancouver

Hemmat A, Vahidi A, Shidani A, Sanian MV, Asadollahzadeh H, Parast AY, et al. ORCA: Hunting Compositional Failures in Text-to-Image Diffusion. https://omanscience.com/en/articles/orca-hunting-compositional-failures-in-text-to-image-diffusion

IEEE

A. Hemmat, A. Vahidi, A. Shidani, M. V. Sanian, H. Asadollahzadeh, A. Y. Parast, and M. Lotfollahi, "ORCA: Hunting Compositional Failures in Text-to-Image Diffusion," https://omanscience.com/en/articles/orca-hunting-compositional-failures-in-text-to-image-diffusion.