Abstract
Discrete action tokenization is central to autoregressive vision-language-action (VLA) models, yet action representations are often evaluated primarily through reconstruction fidelity. We ask which representation properties actually matter for closed-loop control by comparing fixed analytical, data-driven linear, and nonlinear neural representations under a unified tokenization interface. Across rate-distortion analysis, sequence-modeling diagnostics, and 3,500 LIBERO rollouts, representation rankings change with the evaluation criterion. PCA achieves lower nominal reconstruction error than Temporal-DCT, but produces less predictable token sequences and 3.0 percentage points lower mean seen-task success across three policy-training seeds, with the policy ordering reversing in one seed. In a matched seed-42 ablation, an autoencoder further reduces reconstruction error yet does not yield the strongest policy and exhibits greater sensitivity to discrete token perturbations. These findings show that reconstruction fidelity alone cannot reliably select action representations for autoregressive control, motivating joint evaluation of geometric fidelity, sequence predictability, decoder stability, and closed-loop performance.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Yang, Y., He, G., Guan, C., & Liu, H. (2026). Beyond Reconstruction Error: Analytical and Data-Driven Action Tokenization for Autoregressive Vision-Language-Action Models. https://omanscience.com/en/articles/beyond-reconstruction-error-analytical-and-data-driven-action-tokenization-for-autoregressive-vision-language-action-models
MLA 9
Yang, Yuxin, et al. "Beyond Reconstruction Error: Analytical and Data-Driven Action Tokenization for Autoregressive Vision-Language-Action Models." https://omanscience.com/en/articles/beyond-reconstruction-error-analytical-and-data-driven-action-tokenization-for-autoregressive-vision-language-action-models.
Chicago (author–date)
Yang, Yuxin, Gaohan He, Changxue Guan, and Hangming Liu. 2026. "Beyond Reconstruction Error: Analytical and Data-Driven Action Tokenization for Autoregressive Vision-Language-Action Models." https://omanscience.com/en/articles/beyond-reconstruction-error-analytical-and-data-driven-action-tokenization-for-autoregressive-vision-language-action-models.
Harvard
Yang, Y., He, G., Guan, C. and Liu, H. (2026) 'Beyond Reconstruction Error: Analytical and Data-Driven Action Tokenization for Autoregressive Vision-Language-Action Models', Available at: https://omanscience.com/en/articles/beyond-reconstruction-error-analytical-and-data-driven-action-tokenization-for-autoregressive-vision-language-action-models.
Vancouver
Yang Y, He G, Guan C, Liu H. Beyond Reconstruction Error: Analytical and Data-Driven Action Tokenization for Autoregressive Vision-Language-Action Models. https://omanscience.com/en/articles/beyond-reconstruction-error-analytical-and-data-driven-action-tokenization-for-autoregressive-vision-language-action-models
IEEE
Y. Yang, G. He, C. Guan, and H. Liu, "Beyond Reconstruction Error: Analytical and Data-Driven Action Tokenization for Autoregressive Vision-Language-Action Models," https://omanscience.com/en/articles/beyond-reconstruction-error-analytical-and-data-driven-action-tokenization-for-autoregressive-vision-language-action-models.