Abstract
Visual grounding in remote sensing images aims to locate objects described by referring expressions. Most existing methods predict horizontal bounding boxes, which are often inaccurate for objects with arbitrary orientations. To address this limitation, we introduce O$^2$-VG, a family of models for oriented object visual grounding with three complementary designs. Specifically, O$^2$-VG-Trans is a cross-modality transformer for oriented object visual grounding. It establishes a strong discriminative foundation for the model family. Building upon it, O$^2$-VG-Uni predicts universal oriented proposals for possible foreground objects without specific text prompts. It also supports object retrieval through cached proposal embeddings. Using these universal oriented proposals as input prompts, O$^2$-VG-VLM is an autoregressive vision-language model. It generates oriented box token blocks in parallel through multi-token prediction. In addition, we construct DIOR-R-RSVG, a dataset for oriented object visual grounding in remote sensing images. It provides image, expression, and oriented box triplets for training and evaluation. Together, the O$^2$-VG family provides a flexible framework that spans discriminative transformers and generative vision-language models. It achieves superior performance across multiple benchmarks. Code is available at https://github.com/wokaikaixinxin/ai4rs and https://github.com/wokaikaixinxin/Eagle_o2_vg.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Ding, Z., Zhou, Y., Zhao, J., Du, W. L., Li, X., Zhu, H., Yao, R., & El Saddik, A. (2026). A Unified Framework and Dataset for Oriented Object Visual Grounding in Remote Sensing. https://omanscience.com/en/articles/a-unified-framework-and-dataset-for-oriented-object-visual-grounding-in-remote-sensing
MLA 9
Ding, Zeyu, et al. "A Unified Framework and Dataset for Oriented Object Visual Grounding in Remote Sensing." https://omanscience.com/en/articles/a-unified-framework-and-dataset-for-oriented-object-visual-grounding-in-remote-sensing.
Chicago (author–date)
Ding, Zeyu, Yong Zhou, Jiaqi Zhao, Wen-Liang Du, Xixi Li, Hancheng Zhu, Rui Yao, and Abdulmotaleb El Saddik. 2026. "A Unified Framework and Dataset for Oriented Object Visual Grounding in Remote Sensing." https://omanscience.com/en/articles/a-unified-framework-and-dataset-for-oriented-object-visual-grounding-in-remote-sensing.
Harvard
Ding, Z., Zhou, Y., Zhao, J., Du, W. L., Li, X., Zhu, H., Yao, R. and El Saddik, A. (2026) 'A Unified Framework and Dataset for Oriented Object Visual Grounding in Remote Sensing', Available at: https://omanscience.com/en/articles/a-unified-framework-and-dataset-for-oriented-object-visual-grounding-in-remote-sensing.
Vancouver
Ding Z, Zhou Y, Zhao J, Du WL, Li X, Zhu H, et al. A Unified Framework and Dataset for Oriented Object Visual Grounding in Remote Sensing. https://omanscience.com/en/articles/a-unified-framework-and-dataset-for-oriented-object-visual-grounding-in-remote-sensing
IEEE
Z. Ding, Y. Zhou, J. Zhao, W. L. Du, X. Li, H. Zhu, R. Yao, and A. El Saddik, "A Unified Framework and Dataset for Oriented Object Visual Grounding in Remote Sensing," https://omanscience.com/en/articles/a-unified-framework-and-dataset-for-oriented-object-visual-grounding-in-remote-sensing.