الباحثون

Rongxue Li

المنشورات 2

نسخة أولية وصول مفتوح

SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces

Rongxue Li, Meng Yang, Yiru Mao وآخرون · 2026

Spatial coding agents significantly improve spatial reasoning in Multimodal Large Language Models (MLLMs) by using external tools to generate verified execution traces. However, this paradigm inherently suffers from prohibitive inference-time overhead and external dependencies. In this paper, we explore whether an MLLM …

نسخة أولية وصول مفتوح

IronViT: Toward Efficient Generalist Visual Representation Learning

Jiaxi Huang, Yueqi Hu, Xin Zhu وآخرون · 2026

A generalist vision encoder must capture semantic, spatial, language-aligned, and action-relevant cues within a unified representation, yet softmax attention underlying today's most capable visual backbones becomes prohibitively expensive at high resolution. A natural attempt to address both challenges is to distill mu …

المؤلفون المشاركون