Abstract

This paper studies the problem of learning disentangled representations of objects and their attributes from raw, unstructured image data. Slot-based methods have shown considerable success in unsupervised learning of object representations from images. Block-slot attention-based methods extend this framework to attribute representations by assuming a uniform factorization of object representations into attributes, which may be suboptimal and consequently limit the quality of the learned representations. We therefore investigate a framework for jointly discovering object and attribute representations. Our key contribution is leveraging the Linear Representation Hypothesis (LRH), which postulates that composable concepts can be represented as linearly additive subspaces in slot representations. Based on this insight, we propose a probabilistic model connecting images, slots (objects), and blocks (attributes). We present an architecture that leverages block attention to connect attribute representations to slots and incorporates LRH in both object and attribute representation spaces. This architecture effectively optimizes the Evidence Lower Bound (ELBO) of the proposed graphical model. Our experiments demonstrate (i) effective discovery of disentangled object and attribute representations, (ii) empirical evidence for LRH in slot space, and (iii) the ability to perform image editing owing to the disentangled and interpretable nature of the learned representations. Our experiments on multiple datasets demonstrate improvements in DCI scores over state-of-the-art methods.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Gandhi, S., Giri, U., Subramanium, V., Paul, R., & Singla, P. (2026). LinSlot: Exploiting Linear Representation hypothesis for unsupervised attribute discovery from slot based object representation. https://omanscience.com/en/articles/linslot-exploiting-linear-representation-hypothesis-for-unsupervised-attribute-discovery-from-slot-based-object-representation

MLA 9

Gandhi, Sanket, et al. "LinSlot: Exploiting Linear Representation hypothesis for unsupervised attribute discovery from slot based object representation." https://omanscience.com/en/articles/linslot-exploiting-linear-representation-hypothesis-for-unsupervised-attribute-discovery-from-slot-based-object-representation.

Chicago (author–date)

Gandhi, Sanket, Utkarsh Giri, Varun Subramanium, Rohan Paul, and Parag Singla. 2026. "LinSlot: Exploiting Linear Representation hypothesis for unsupervised attribute discovery from slot based object representation." https://omanscience.com/en/articles/linslot-exploiting-linear-representation-hypothesis-for-unsupervised-attribute-discovery-from-slot-based-object-representation.

Harvard

Gandhi, S., Giri, U., Subramanium, V., Paul, R. and Singla, P. (2026) 'LinSlot: Exploiting Linear Representation hypothesis for unsupervised attribute discovery from slot based object representation', Available at: https://omanscience.com/en/articles/linslot-exploiting-linear-representation-hypothesis-for-unsupervised-attribute-discovery-from-slot-based-object-representation.

Vancouver

Gandhi S, Giri U, Subramanium V, Paul R, Singla P. LinSlot: Exploiting Linear Representation hypothesis for unsupervised attribute discovery from slot based object representation. https://omanscience.com/en/articles/linslot-exploiting-linear-representation-hypothesis-for-unsupervised-attribute-discovery-from-slot-based-object-representation

IEEE

S. Gandhi, U. Giri, V. Subramanium, R. Paul, and P. Singla, "LinSlot: Exploiting Linear Representation hypothesis for unsupervised attribute discovery from slot based object representation," https://omanscience.com/en/articles/linslot-exploiting-linear-representation-hypothesis-for-unsupervised-attribute-discovery-from-slot-based-object-representation.