الباحثون

Lulu Hu

المنشورات 2

نسخة أولية وصول مفتوح

V-CoLA: Vision Token Compression with Linear Attention

Hao Jiang, Yiru Mao, Tianpeng Bu وآخرون · 2026

Vision-language models (VLMs) have demonstrated impressive capabilities but suffer from substantial computational overhead, as vision tokens dominate the input sequence. This motivates vision token compression as a key direction to alleviate the burden. However, with the emergence of hybrid architectures incorporating …

نسخة أولية وصول مفتوح

SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces

Rongxue Li, Meng Yang, Yiru Mao وآخرون · 2026

Spatial coding agents significantly improve spatial reasoning in Multimodal Large Language Models (MLLMs) by using external tools to generate verified execution traces. However, this paradigm inherently suffers from prohibitive inference-time overhead and external dependencies. In this paper, we explore whether an MLLM …

المؤلفون المشاركون