Abstract
Low-bit vision-language-action inference must reduce observation-to-action latency while preserving robot behavior. We present FoldQuantVLA, a post-training quantization framework that carries a consistent activation representation through calibration, weight rounding, and native integer execution. It combines channel scaling and block Hadamard transforms with dynamic per-token quantization, without policy retraining. Custom TensorRT plugins execute projections in both the language backbone and iterative action expert with four-bit weights and activations (W4A4) on Ada GPUs and Jetson AGX Orin. Evaluation spans LIBERO, SimplerEnv, and two robot platforms. Across three GR00T checkpoints and $π_{0.5}$, W4A4 achieves $1.20$ to $1.33\times$ speedups over floating-point TensorRT on Orin and $1.25$ to $1.52\times$ on desktop. Retaining language attention-output and feed-forward down projections at eight bits (W8A8) improves held-out action fidelity on all four checkpoints. Across four real-robot tasks, this configuration raises observed GR00T N1.7 success from $80.0\%$ with uniform W4A4 to $92.5\%$ over 80 trials per configuration, with a measured additional Orin latency of 1 ms.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Ho, H. T., Nguyen, K. D., Nguyen, Q. D., Duong, T. Q., Le, N., Guo, M., Ngo, V. A., & Le, A. T. (2026). FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding. https://omanscience.com/en/articles/foldquantvla-native-low-bit-quantization-of-vision-language-action-models-via-consistent-folding
MLA 9
Ho, Hung T., et al. "FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding." https://omanscience.com/en/articles/foldquantvla-native-low-bit-quantization-of-vision-language-action-models-via-consistent-folding.
Chicago (author–date)
Ho, Hung T., Khanh D. Nguyen, Quang D. Nguyen, Thanh Q. Duong, Ngan Le, Meng Guo, Vien A. Ngo, and An T. Le. 2026. "FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding." https://omanscience.com/en/articles/foldquantvla-native-low-bit-quantization-of-vision-language-action-models-via-consistent-folding.
Harvard
Ho, H. T., Nguyen, K. D., Nguyen, Q. D., Duong, T. Q., Le, N., Guo, M., Ngo, V. A. and Le, A. T. (2026) 'FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding', Available at: https://omanscience.com/en/articles/foldquantvla-native-low-bit-quantization-of-vision-language-action-models-via-consistent-folding.
Vancouver
Ho HT, Nguyen KD, Nguyen QD, Duong TQ, Le N, Guo M, et al. FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding. https://omanscience.com/en/articles/foldquantvla-native-low-bit-quantization-of-vision-language-action-models-via-consistent-folding
IEEE
H. T. Ho, K. D. Nguyen, Q. D. Nguyen, T. Q. Duong, N. Le, M. Guo, V. A. Ngo, and A. T. Le, "FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding," https://omanscience.com/en/articles/foldquantvla-native-low-bit-quantization-of-vision-language-action-models-via-consistent-folding.