Abstract
Tokens have become a unified interface for multimodal foundation models, making visual-token communication a natural paradigm for efficient image delivery. However, existing methods typically rely on static policies that cannot jointly adapt to image content and channel conditions. Moreover, their token-level utility objectives do not necessarily translate into improved image reconstruction quality. In this paper, we propose AdapToC, an adaptive, reconstruction-oriented visual-token communication framework. At the transmitter, an adaptive selector jointly models image content, channel state, and communication budget to perform instance-wise resource allocation. Rather than using a fixed token rate and protection policy, it dynamically determines how many tokens should be transmitted and assigns different protection levels according to token importance and current channel conditions. At the receiver, an adaptive MaskGIT receiver incorporates channel reliability into contextual token modeling. It distinguishes tokens with different reliability levels, preserves high-confidence observations, corrects potentially corrupted tokens, and iteratively reconstructs missing content from the received evidence and global visual context. By co-designing token quantity, unequal protection, and reliability-aware recovery for image-level reconstruction quality, AdapToC achieves a peak mean PSNR gain of 4.20 dB over the strongest static baseline under matched communication costs and state-of-the-art performance among the evaluated visual-token communication methods.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Liu, Z., Zhao, Z., Wu, C., Yang, Z., & Zhang, Z. (2026). Beyond Token Accuracy: Prioritizing What Matters for Visual Reconstruction. https://omanscience.com/en/articles/beyond-token-accuracy-prioritizing-what-matters-for-visual-reconstruction
MLA 9
Liu, Zhicheng, et al. "Beyond Token Accuracy: Prioritizing What Matters for Visual Reconstruction." https://omanscience.com/en/articles/beyond-token-accuracy-prioritizing-what-matters-for-visual-reconstruction.
Chicago (author–date)
Liu, Zhicheng, Zhouxiang Zhao, Chenliang Wu, Zhaohui Yang, and Zhaoyang Zhang. 2026. "Beyond Token Accuracy: Prioritizing What Matters for Visual Reconstruction." https://omanscience.com/en/articles/beyond-token-accuracy-prioritizing-what-matters-for-visual-reconstruction.
Harvard
Liu, Z., Zhao, Z., Wu, C., Yang, Z. and Zhang, Z. (2026) 'Beyond Token Accuracy: Prioritizing What Matters for Visual Reconstruction', Available at: https://omanscience.com/en/articles/beyond-token-accuracy-prioritizing-what-matters-for-visual-reconstruction.
Vancouver
Liu Z, Zhao Z, Wu C, Yang Z, Zhang Z. Beyond Token Accuracy: Prioritizing What Matters for Visual Reconstruction. https://omanscience.com/en/articles/beyond-token-accuracy-prioritizing-what-matters-for-visual-reconstruction
IEEE
Z. Liu, Z. Zhao, C. Wu, Z. Yang, and Z. Zhang, "Beyond Token Accuracy: Prioritizing What Matters for Visual Reconstruction," https://omanscience.com/en/articles/beyond-token-accuracy-prioritizing-what-matters-for-visual-reconstruction.