Abstract
Action tokenization converts continuous robot actions into discrete symbols that can be modeled autoregressively. However, existing tokenizer-based policies typically ignore the tokenizer's learned latent code structure: after tokenization, the policy treats tokens as unrelated class indices and learns a new classifier from scratch. We show that this discarded structure is valuable. We introduce Codebook-Aligned Prediction (CAP), a method that directly reuses the tokenizer's code vectors as policy class prototypes while leaving the tokenizer and policy backbone otherwise unchanged. Across four quantizer families, three simulation benchmarks, and two real-robot tasks, CAP consistently improves task success over standard token classification heads while holding the tokenizer (and therefore its reconstruction quality) fixed. Our analysis further shows that these gains are not explained by higher token accuracy or changes in the policy head alone. Instead, reusing the tokenizer codebook provides the policy with valuable information about the tokenizer's learned latent structure across tokens, making token prediction errors more benign in action space and improving the representations learned by the policy backbone. These results suggest that action tokenizers learn useful action-aware latent structure beyond discrete targets that should be preserved when training downstream policies.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Chen, H., Ji, J., Wheeler, S., Stocking, K. C., & Walter, M. (2026). CAP: Codebook-Aligned Prediction for Tokenized Robot Policies. https://omanscience.com/en/articles/cap-codebook-aligned-prediction-for-tokenized-robot-policies
MLA 9
Chen, Haoran, et al. "CAP: Codebook-Aligned Prediction for Tokenized Robot Policies." https://omanscience.com/en/articles/cap-codebook-aligned-prediction-for-tokenized-robot-policies.
Chicago (author–date)
Chen, Haoran, Jingtian Ji, Samuel Wheeler, Kaylene Caswell Stocking, and Matthew Walter. 2026. "CAP: Codebook-Aligned Prediction for Tokenized Robot Policies." https://omanscience.com/en/articles/cap-codebook-aligned-prediction-for-tokenized-robot-policies.
Harvard
Chen, H., Ji, J., Wheeler, S., Stocking, K. C. and Walter, M. (2026) 'CAP: Codebook-Aligned Prediction for Tokenized Robot Policies', Available at: https://omanscience.com/en/articles/cap-codebook-aligned-prediction-for-tokenized-robot-policies.
Vancouver
Chen H, Ji J, Wheeler S, Stocking KC, Walter M. CAP: Codebook-Aligned Prediction for Tokenized Robot Policies. https://omanscience.com/en/articles/cap-codebook-aligned-prediction-for-tokenized-robot-policies
IEEE
H. Chen, J. Ji, S. Wheeler, K. C. Stocking, and M. Walter, "CAP: Codebook-Aligned Prediction for Tokenized Robot Policies," https://omanscience.com/en/articles/cap-codebook-aligned-prediction-for-tokenized-robot-policies.