Abstract

Discrete action tokenization is central to autoregressive vision-language-action (VLA) models, yet action representations are often evaluated primarily through reconstruction fidelity. We ask which representation properties actually matter for closed-loop control by comparing fixed analytical, data-driven linear, and nonlinear neural representations under a unified tokenization interface. Across rate-distortion analysis, sequence-modeling diagnostics, and 3,500 LIBERO rollouts, representation rankings change with the evaluation criterion. PCA achieves lower nominal reconstruction error than Temporal-DCT, but produces less predictable token sequences and 3.0 percentage points lower mean seen-task success across three policy-training seeds, with the policy ordering reversing in one seed. In a matched seed-42 ablation, an autoencoder further reduces reconstruction error yet does not yield the strongest policy and exhibits greater sensitivity to discrete token perturbations. These findings show that reconstruction fidelity alone cannot reliably select action representations for autoregressive control, motivating joint evaluation of geometric fidelity, sequence predictability, decoder stability, and closed-loop performance.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Yang, Y., He, G., Guan, C., & Liu, H. (2026). Beyond Reconstruction Error: Analytical and Data-Driven Action Tokenization for Autoregressive Vision-Language-Action Models. https://omanscience.com/en/articles/beyond-reconstruction-error-analytical-and-data-driven-action-tokenization-for-autoregressive-vision-language-action-models

MLA 9

Yang, Yuxin, et al. "Beyond Reconstruction Error: Analytical and Data-Driven Action Tokenization for Autoregressive Vision-Language-Action Models." https://omanscience.com/en/articles/beyond-reconstruction-error-analytical-and-data-driven-action-tokenization-for-autoregressive-vision-language-action-models.

Chicago (author–date)

Yang, Yuxin, Gaohan He, Changxue Guan, and Hangming Liu. 2026. "Beyond Reconstruction Error: Analytical and Data-Driven Action Tokenization for Autoregressive Vision-Language-Action Models." https://omanscience.com/en/articles/beyond-reconstruction-error-analytical-and-data-driven-action-tokenization-for-autoregressive-vision-language-action-models.

Harvard

Yang, Y., He, G., Guan, C. and Liu, H. (2026) 'Beyond Reconstruction Error: Analytical and Data-Driven Action Tokenization for Autoregressive Vision-Language-Action Models', Available at: https://omanscience.com/en/articles/beyond-reconstruction-error-analytical-and-data-driven-action-tokenization-for-autoregressive-vision-language-action-models.

Vancouver

Yang Y, He G, Guan C, Liu H. Beyond Reconstruction Error: Analytical and Data-Driven Action Tokenization for Autoregressive Vision-Language-Action Models. https://omanscience.com/en/articles/beyond-reconstruction-error-analytical-and-data-driven-action-tokenization-for-autoregressive-vision-language-action-models

IEEE

Y. Yang, G. He, C. Guan, and H. Liu, "Beyond Reconstruction Error: Analytical and Data-Driven Action Tokenization for Autoregressive Vision-Language-Action Models," https://omanscience.com/en/articles/beyond-reconstruction-error-analytical-and-data-driven-action-tokenization-for-autoregressive-vision-language-action-models.