Abstract

Tokens have become a unified interface for multimodal foundation models, making visual-token communication a natural paradigm for efficient image delivery. However, existing methods typically rely on static policies that cannot jointly adapt to image content and channel conditions. Moreover, their token-level utility objectives do not necessarily translate into improved image reconstruction quality. In this paper, we propose AdapToC, an adaptive, reconstruction-oriented visual-token communication framework. At the transmitter, an adaptive selector jointly models image content, channel state, and communication budget to perform instance-wise resource allocation. Rather than using a fixed token rate and protection policy, it dynamically determines how many tokens should be transmitted and assigns different protection levels according to token importance and current channel conditions. At the receiver, an adaptive MaskGIT receiver incorporates channel reliability into contextual token modeling. It distinguishes tokens with different reliability levels, preserves high-confidence observations, corrects potentially corrupted tokens, and iteratively reconstructs missing content from the received evidence and global visual context. By co-designing token quantity, unequal protection, and reliability-aware recovery for image-level reconstruction quality, AdapToC achieves a peak mean PSNR gain of 4.20 dB over the strongest static baseline under matched communication costs and state-of-the-art performance among the evaluated visual-token communication methods.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Liu, Z., Zhao, Z., Wu, C., Yang, Z., & Zhang, Z. (2026). Beyond Token Accuracy: Prioritizing What Matters for Visual Reconstruction. https://omanscience.com/en/articles/beyond-token-accuracy-prioritizing-what-matters-for-visual-reconstruction

MLA 9

Liu, Zhicheng, et al. "Beyond Token Accuracy: Prioritizing What Matters for Visual Reconstruction." https://omanscience.com/en/articles/beyond-token-accuracy-prioritizing-what-matters-for-visual-reconstruction.

Chicago (author–date)

Liu, Zhicheng, Zhouxiang Zhao, Chenliang Wu, Zhaohui Yang, and Zhaoyang Zhang. 2026. "Beyond Token Accuracy: Prioritizing What Matters for Visual Reconstruction." https://omanscience.com/en/articles/beyond-token-accuracy-prioritizing-what-matters-for-visual-reconstruction.

Harvard

Liu, Z., Zhao, Z., Wu, C., Yang, Z. and Zhang, Z. (2026) 'Beyond Token Accuracy: Prioritizing What Matters for Visual Reconstruction', Available at: https://omanscience.com/en/articles/beyond-token-accuracy-prioritizing-what-matters-for-visual-reconstruction.

Vancouver

Liu Z, Zhao Z, Wu C, Yang Z, Zhang Z. Beyond Token Accuracy: Prioritizing What Matters for Visual Reconstruction. https://omanscience.com/en/articles/beyond-token-accuracy-prioritizing-what-matters-for-visual-reconstruction

IEEE

Z. Liu, Z. Zhao, C. Wu, Z. Yang, and Z. Zhang, "Beyond Token Accuracy: Prioritizing What Matters for Visual Reconstruction," https://omanscience.com/en/articles/beyond-token-accuracy-prioritizing-what-matters-for-visual-reconstruction.