الملخص

Vision Transformers (ViTs) achieve strong performance on image recognition and mobile vision applications, but their high-dimensional linear projections and attention computations still impose substantial storage and inference costs. Extremely low-bit quantization is a promising solution, yet ViTs often suffer severe accuracy degradation because conventional real-valued scalar codebooks are poorly matched to the directional geometry of Transformer projections. We present RPFQ-ViT, a Rotated Phase-Frame Quantization method that quantizes paired channels in two-dimensional phase planes, enabling low-bit codes to better preserve projection directions while recovering magnitude with lightweight scaling. RPFQ-ViT serves as a drop-in QAT replacement for nn.Linear and does not modify the standard real-valued attention, normalization, or activation computation graph. On ImageNet-1K, RPFQ-ViT-B/16 reaches 79.33% Top-1 / 94.48% Top-5 under W2/A4, Swin-T reaches 79.30% Top-1 / 94.79% Top-5 under W2/A8, and DeiT-S reaches 77.41% Top-1 / 93.11% Top-5 under W2/A8. Ablations, phase-geometry analysis, and direction-preservation metrics show that channel pairing, learnable rotation, phase-anchor learning, and residual phase refinement each improve quantization quality. We further deploy RPFQ-ViT image-classification models on native iOS and Android runtime stacks; with 2-bit packed weights, model size shrinks by roughly $5.4$-$7.1\times$ relative to FP32 and end-to-end on-device latency drops by $1.4$-$1.6\times$. All ImageNet results trained in our codebase use a matched 300-epoch recipe and are reported as mean accuracies over three independent runs. These results show that RPFQ-ViT provides a favorable trade-off among accuracy, compression, and practical mobile deployment for extremely low-bit ViTs.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Fan, M., Huang, B., Pan, J., Yuan, X., Cong, P., Tan, Z., & Yang, T. (2026). RPFQ-ViT: Rotated Phase-Frame Quantization for Extremely Low-Bit Weights in Vision Transformers. https://omanscience.com/ar/articles/rpfq-vit-rotated-phase-frame-quantization-for-extremely-low-bit-weights-in-vision-transformers

MLA 9

Fan, Mengyuan, et al. "RPFQ-ViT: Rotated Phase-Frame Quantization for Extremely Low-Bit Weights in Vision Transformers." https://omanscience.com/ar/articles/rpfq-vit-rotated-phase-frame-quantization-for-extremely-low-bit-weights-in-vision-transformers.

شيكاغو (المؤلف–التاريخ)

Fan, Mengyuan, Bokai Huang, JiaMing Pan, Xiaokun Yuan, Peizhuang Cong, Zhewen Tan, and Tong Yang. 2026. "RPFQ-ViT: Rotated Phase-Frame Quantization for Extremely Low-Bit Weights in Vision Transformers." https://omanscience.com/ar/articles/rpfq-vit-rotated-phase-frame-quantization-for-extremely-low-bit-weights-in-vision-transformers.

هارفارد

Fan, M., Huang, B., Pan, J., Yuan, X., Cong, P., Tan, Z. and Yang, T. (2026) 'RPFQ-ViT: Rotated Phase-Frame Quantization for Extremely Low-Bit Weights in Vision Transformers', Available at: https://omanscience.com/ar/articles/rpfq-vit-rotated-phase-frame-quantization-for-extremely-low-bit-weights-in-vision-transformers.

فانكوفر

Fan M, Huang B, Pan J, Yuan X, Cong P, Tan Z, et al. RPFQ-ViT: Rotated Phase-Frame Quantization for Extremely Low-Bit Weights in Vision Transformers. https://omanscience.com/ar/articles/rpfq-vit-rotated-phase-frame-quantization-for-extremely-low-bit-weights-in-vision-transformers

IEEE

M. Fan, B. Huang, J. Pan, X. Yuan, P. Cong, Z. Tan, and T. Yang, "RPFQ-ViT: Rotated Phase-Frame Quantization for Extremely Low-Bit Weights in Vision Transformers," https://omanscience.com/ar/articles/rpfq-vit-rotated-phase-frame-quantization-for-extremely-low-bit-weights-in-vision-transformers.