الملخص
Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply an intrinsic left-to-right generation order. This distinction makes bidirectional diffusion a natural fit, allowing spatial hypotheses to emerge in parallel and be jointly refined through iterative denoising. We introduce GroundAnything, a 4B-parameter grounding foundation model that reconciles fast parallel decoding with precise localization through blockwise denoising. Training combines grounding pretraining from public datasets and dedicated data engines, direct AR-to-diffusion conversion with joint AR and diffusion objectives, supervised fine-tuning, and GRPO-based reinforcement post-training. Across 30 grounding benchmarks, our autoregressive variant, GroundAnything-VLM, establishes a new overall state of the art among similarly sized models at 72.42%, remaining competitive with GPT-6 Astra (71.35%). With entropy-guided decoding, GroundAnything also surpasses the prior state of the art at this scale, averaging 61.75% versus 53.32% for the fast MTP-based LocateAnything model. We further explore decoding strategies, showing that an optional self-speculative mode achieves a $4.51\times$ speedup over the AR counterpart with a 0.74 percentage-point drop in COCO F1mIoU. Infrastructure experiments show that progressive inference optimizations translate parallel decoding into practical speedups. These support efficient visual grounding in latency-sensitive real-world systems.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Yu, Q., Fan, L., Ping, B., Ding, X., Song, Z., Niu, J., Wang, K., Chen, T., Chen, Y., He, M., Wang, Y., Huang, J., Zhang, H., Chen, M., Li, H., Song, W., Wu, R., Liu, X., Liu, S., Zhou, S., Luo, P., & Huang, S. (2026). GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed. https://omanscience.com/ar/articles/groundanything-reconciling-parallel-decoding-with-precise-visual-grounding-at-flash-speed
MLA 9
Yu, Qize, et al. "GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed." https://omanscience.com/ar/articles/groundanything-reconciling-parallel-decoding-with-precise-visual-grounding-at-flash-speed.
شيكاغو (المؤلف–التاريخ)
Yu, Qize, Lianrui Fan, Bowen Ping, Xini Ding, Zetian Song, Junbo Niu, Kaixuan Wang, Tianxing Chen, Yue Chen, Minghua He, Yuran Wang, Jie Huang, Haojun Zhang, Min Chen, Hao Li, Wenxuan Song, Ruihai Wu, Xianming Liu, Shilong Liu, Shuchang Zhou, Ping Luo, and Shiyu Huang. 2026. "GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed." https://omanscience.com/ar/articles/groundanything-reconciling-parallel-decoding-with-precise-visual-grounding-at-flash-speed.
هارفارد
Yu, Q., Fan, L., Ping, B., Ding, X., Song, Z., Niu, J., Wang, K., Chen, T., Chen, Y., He, M., Wang, Y., Huang, J., Zhang, H., Chen, M., Li, H., Song, W., Wu, R., Liu, X., Liu, S., Zhou, S., Luo, P. and Huang, S. (2026) 'GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed', Available at: https://omanscience.com/ar/articles/groundanything-reconciling-parallel-decoding-with-precise-visual-grounding-at-flash-speed.
فانكوفر
Yu Q, Fan L, Ping B, Ding X, Song Z, Niu J, et al. GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed. https://omanscience.com/ar/articles/groundanything-reconciling-parallel-decoding-with-precise-visual-grounding-at-flash-speed
IEEE
Q. Yu, L. Fan, B. Ping, X. Ding, Z. Song, J. Niu, K. Wang, T. Chen, Y. Chen, M. He, Y. Wang, J. Huang, H. Zhang, M. Chen, H. Li, W. Song, R. Wu, X. Liu, S. Liu, S. Zhou, P. Luo, and S. Huang, "GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed," https://omanscience.com/ar/articles/groundanything-reconciling-parallel-decoding-with-precise-visual-grounding-at-flash-speed.