الباحثون

Shaoyi Huang

المنشورات 3

نسخة أولية وصول مفتوح

Imagine the Future, Internalize the Gist: Efficient VLA Reasoning via Internalized Spatiotemporal Imagination

Shenglan Li, Zhendong Mi, Hengyi Zhu وآخرون · 2026

Vision-language-action (VLA) models increasingly incorporate intermediate reasoning to improve robotic manipulation, yet existing approaches primarily reason about observed states without explicitly anticipating future scene evolution. Extending such reasoning to explicit future rollouts at every inference step, howeve …

نسخة أولية وصول مفتوح

From Patching to Pruning Visual Computation in Vision Language Models

Vision language models (VLMs) incur substantial inference cost because every visual token is processed by the attention and MLP projections of every decoder layer, even when token-specific visual computation is unnecessary at many depths. We introduce Patch-to-Prune (P2P), inspired by Mechanistic Interpretability, a tr …

نسخة أولية وصول مفتوح

SpectralCache: Accelerating Diffusion-Based World Models via Spectral Feature Caching

Zhendong Mi, Pu Zhao, Ziyu Hu وآخرون · 2026

Diffusion-based world models enable high-quality interactive environment generation but suffer from substantial inference overhead due to repeated Transformer evaluations during denoising. Existing caching methods mainly exploit temporal redundancy at the feature or token level, leaving the underlying mathematical stru …

المؤلفون المشاركون