Abstract
Open-vocabulary detection (OVD) recognizes categories unseen during training through textual category queries, yet achieving strong generalization with real-time efficiency remains challenging. Beyond vocabulary scaling, zero-shot generalization may benefit from reusable visual--semantic cues learned from seen data, including attributes, actions, states, and contextual relations. Existing real-time OVD methods primarily emphasize vocabulary coverage and efficient region/query--text matching; under strict efficiency constraints, compact detectors may struggle to absorb rich instance semantics and scene context. We propose RT-DETR-World, a compact DETR-style detector that transfers the rich semantics conveyed by descriptions during training while retaining lightweight query--text matching at inference. We construct GroundingCapv2 with three levels of supervision: category names for standard OVD, object descriptions conveying instance-level semantics, and image descriptions conveying object relations and scene context. These descriptions serve only as training-time semantic supervision. To help the compact detector absorb these semantics, we propose Dual-Path Description Alignment (DDA), combining a deployment-consistent MiniLM pathway with a training-only LLM teacher. MiniLM provides query--category supervision and object-description alignment, while offline teacher features supervise matched queries and global visual representations at the object and image levels, respectively. All teacher features are precomputed, and the teacher-side modules are removed after training. We further propose Relation-Aware Negative Relaxation (RNR), which uses teacher-derived semantic similarities to relax related negatives while preserving exact positives. Experiments demonstrate competitive zero-shot accuracy and a favorable accuracy--efficiency trade-off. The code will be released.
Keywords
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Zhang, Y., Zhao, Z., Cheng, J., Wang, S., Guo, N., Han, R., & Wan, L. (2026). RT-DETR-World: Transferring Rich LLM Semantics to Real-Time Open-Vocabulary Detection. https://omanscience.com/en/articles/rt-detr-world-transferring-rich-llm-semantics-to-real-time-open-vocabulary-detection
MLA 9
Zhang, Yupeng, et al. "RT-DETR-World: Transferring Rich LLM Semantics to Real-Time Open-Vocabulary Detection." https://omanscience.com/en/articles/rt-detr-world-transferring-rich-llm-semantics-to-real-time-open-vocabulary-detection.
Chicago (author–date)
Zhang, Yupeng, Ziyi Zhao, Juntao Cheng, Sheng Wang, Ningnan Guo, Ruize Han, and Liang Wan. 2026. "RT-DETR-World: Transferring Rich LLM Semantics to Real-Time Open-Vocabulary Detection." https://omanscience.com/en/articles/rt-detr-world-transferring-rich-llm-semantics-to-real-time-open-vocabulary-detection.
Harvard
Zhang, Y., Zhao, Z., Cheng, J., Wang, S., Guo, N., Han, R. and Wan, L. (2026) 'RT-DETR-World: Transferring Rich LLM Semantics to Real-Time Open-Vocabulary Detection', Available at: https://omanscience.com/en/articles/rt-detr-world-transferring-rich-llm-semantics-to-real-time-open-vocabulary-detection.
Vancouver
Zhang Y, Zhao Z, Cheng J, Wang S, Guo N, Han R, et al. RT-DETR-World: Transferring Rich LLM Semantics to Real-Time Open-Vocabulary Detection. https://omanscience.com/en/articles/rt-detr-world-transferring-rich-llm-semantics-to-real-time-open-vocabulary-detection
IEEE
Y. Zhang, Z. Zhao, J. Cheng, S. Wang, N. Guo, R. Han, and L. Wan, "RT-DETR-World: Transferring Rich LLM Semantics to Real-Time Open-Vocabulary Detection," https://omanscience.com/en/articles/rt-detr-world-transferring-rich-llm-semantics-to-real-time-open-vocabulary-detection.