Abstract

Multi-task manipulation policies differ in architecture, scale, and pretrained priors all at once, so published comparisons cannot attribute performance to any single component. We argue that most of it comes from the visual representation, and that parameter scale and generative priors are largely incidental. To test this, we build CoRP (Compressed Representation Policy), a deliberately compact policy (48.9M parameters, no vision-language model and no video-generative prior) that factorizes into a representation extractor and a flow-matching action generator. It reaches 97.0% on LIBERO and 75.78%/73.36% on RoboTwin 2.0 Clean/Randomized, matching systems 40.9-163.6x larger. Holding the action generator fixed, we then vary one extractor property at a time. Pretrained initialization is decisive: a random ViT-S/14 drops to 78.1% and an ImageNet ResNet-34 to 74.5% on LIBERO. Pretraining alone is not enough, as freezing the encoder costs 19.8 points. Compression matters as much: resampling each view to 48 tokens beats passing all patch tokens (97.0% vs 83.2%), and a variational information bottleneck over those tokens is worse than a hard token budget, cutting LIBERO-Goal from 95.8% to 33.0% by suppressing the instruction-dependent token selection the policy relies on. Language conditioning contributes only where the observation leaves the goal ambiguous (LIBERO-Goal: 9.2% to 95.8%), while on RoboTwin 2.0, where observations are unambiguous, removing it slightly improves success. Therefore, we argue that a compact policy works when its representation is pretrained, task-adapted, and compressed. Project page: https://corp-policy.github.io/

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Chen, N., Yang, R., Tang, J., Liu, S., & Wang, Y. (2026). Compact Robot Policies Need Fine-Grained Visual Representations. https://omanscience.com/en/articles/compact-robot-policies-need-fine-grained-visual-representations

MLA 9

Chen, Nanhe, et al. "Compact Robot Policies Need Fine-Grained Visual Representations." https://omanscience.com/en/articles/compact-robot-policies-need-fine-grained-visual-representations.

Chicago (author–date)

Chen, Nanhe, Runqiu Yang, Jiawei Tang, Sichao Liu, and Yuquan Wang. 2026. "Compact Robot Policies Need Fine-Grained Visual Representations." https://omanscience.com/en/articles/compact-robot-policies-need-fine-grained-visual-representations.

Harvard

Chen, N., Yang, R., Tang, J., Liu, S. and Wang, Y. (2026) 'Compact Robot Policies Need Fine-Grained Visual Representations', Available at: https://omanscience.com/en/articles/compact-robot-policies-need-fine-grained-visual-representations.

Vancouver

Chen N, Yang R, Tang J, Liu S, Wang Y. Compact Robot Policies Need Fine-Grained Visual Representations. https://omanscience.com/en/articles/compact-robot-policies-need-fine-grained-visual-representations

IEEE

N. Chen, R. Yang, J. Tang, S. Liu, and Y. Wang, "Compact Robot Policies Need Fine-Grained Visual Representations," https://omanscience.com/en/articles/compact-robot-policies-need-fine-grained-visual-representations.