Abstract

Visual grounding in remote sensing images aims to locate objects described by referring expressions. Most existing methods predict horizontal bounding boxes, which are often inaccurate for objects with arbitrary orientations. To address this limitation, we introduce O$^2$-VG, a family of models for oriented object visual grounding with three complementary designs. Specifically, O$^2$-VG-Trans is a cross-modality transformer for oriented object visual grounding. It establishes a strong discriminative foundation for the model family. Building upon it, O$^2$-VG-Uni predicts universal oriented proposals for possible foreground objects without specific text prompts. It also supports object retrieval through cached proposal embeddings. Using these universal oriented proposals as input prompts, O$^2$-VG-VLM is an autoregressive vision-language model. It generates oriented box token blocks in parallel through multi-token prediction. In addition, we construct DIOR-R-RSVG, a dataset for oriented object visual grounding in remote sensing images. It provides image, expression, and oriented box triplets for training and evaluation. Together, the O$^2$-VG family provides a flexible framework that spans discriminative transformers and generative vision-language models. It achieves superior performance across multiple benchmarks. Code is available at https://github.com/wokaikaixinxin/ai4rs and https://github.com/wokaikaixinxin/Eagle_o2_vg.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Ding, Z., Zhou, Y., Zhao, J., Du, W. L., Li, X., Zhu, H., Yao, R., & El Saddik, A. (2026). A Unified Framework and Dataset for Oriented Object Visual Grounding in Remote Sensing. https://omanscience.com/en/articles/a-unified-framework-and-dataset-for-oriented-object-visual-grounding-in-remote-sensing

MLA 9

Ding, Zeyu, et al. "A Unified Framework and Dataset for Oriented Object Visual Grounding in Remote Sensing." https://omanscience.com/en/articles/a-unified-framework-and-dataset-for-oriented-object-visual-grounding-in-remote-sensing.

Chicago (author–date)

Ding, Zeyu, Yong Zhou, Jiaqi Zhao, Wen-Liang Du, Xixi Li, Hancheng Zhu, Rui Yao, and Abdulmotaleb El Saddik. 2026. "A Unified Framework and Dataset for Oriented Object Visual Grounding in Remote Sensing." https://omanscience.com/en/articles/a-unified-framework-and-dataset-for-oriented-object-visual-grounding-in-remote-sensing.

Harvard

Ding, Z., Zhou, Y., Zhao, J., Du, W. L., Li, X., Zhu, H., Yao, R. and El Saddik, A. (2026) 'A Unified Framework and Dataset for Oriented Object Visual Grounding in Remote Sensing', Available at: https://omanscience.com/en/articles/a-unified-framework-and-dataset-for-oriented-object-visual-grounding-in-remote-sensing.

Vancouver

Ding Z, Zhou Y, Zhao J, Du WL, Li X, Zhu H, et al. A Unified Framework and Dataset for Oriented Object Visual Grounding in Remote Sensing. https://omanscience.com/en/articles/a-unified-framework-and-dataset-for-oriented-object-visual-grounding-in-remote-sensing

IEEE

Z. Ding, Y. Zhou, J. Zhao, W. L. Du, X. Li, H. Zhu, R. Yao, and A. El Saddik, "A Unified Framework and Dataset for Oriented Object Visual Grounding in Remote Sensing," https://omanscience.com/en/articles/a-unified-framework-and-dataset-for-oriented-object-visual-grounding-in-remote-sensing.