Abstract
Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI spends tokens on irrelevant content. We introduce FocusVTC, which breaks this trade-off through adaptive resolution while preserving general multimodal capabilities. It combines compressed low-DPI global views with selective region enhancement, integrating enhanced views into ongoing reasoning. We construct 29.4K high-quality Reasoning-Evidence Localization (REL) chain-of-thought examples (REL-CoT) that link reasoning traces to page indices and bounding boxes. Multi-resolution REL supervised fine-tuning (REL-SFT) teaches the model to localize relevant regions, and Group Relative Policy Optimization learns when to enhance resolution and how to use the resulting observations, without a separate continual-pretraining stage. At 72 DPI on RULER v1, FocusVTC scores 87.4 at $2.9\times$ input compression, including tool observations, versus 57.5 for Glyph at $3.0\times$ input compression. It surpasses its text-input backbone on LongBench (56.40 versus 55.86), improves the MRCR macro-average by 13.91 points, and achieves a 51.19 macro-average on VTCBench. The MRCR latency evaluation also shows a $2.79\times$ online end-to-end speedup over Text. General multimodal capabilities are preserved, with MMMU increasing from 65.12 to 66.73 and MME from 2424.02 to 2457.62.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Zhong, F., Qiu, X., Pan, Y., Liu, Y., Gu, S., Xu, B., & Li, G. (2026). FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution. https://omanscience.com/en/articles/focusvtc-efficient-and-high-performance-visual-text-compression-with-adaptive-resolution
MLA 9
Zhong, FangZhi, et al. "FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution." https://omanscience.com/en/articles/focusvtc-efficient-and-high-performance-visual-text-compression-with-adaptive-resolution.
Chicago (author–date)
Zhong, FangZhi, Xuerui Qiu, Yuqi Pan, Ya Liu, Shaowei Gu, Bo Xu, and Guoqi Li. 2026. "FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution." https://omanscience.com/en/articles/focusvtc-efficient-and-high-performance-visual-text-compression-with-adaptive-resolution.
Harvard
Zhong, F., Qiu, X., Pan, Y., Liu, Y., Gu, S., Xu, B. and Li, G. (2026) 'FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution', Available at: https://omanscience.com/en/articles/focusvtc-efficient-and-high-performance-visual-text-compression-with-adaptive-resolution.
Vancouver
Zhong F, Qiu X, Pan Y, Liu Y, Gu S, Xu B, et al. FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution. https://omanscience.com/en/articles/focusvtc-efficient-and-high-performance-visual-text-compression-with-adaptive-resolution
IEEE
F. Zhong, X. Qiu, Y. Pan, Y. Liu, S. Gu, B. Xu, and G. Li, "FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution," https://omanscience.com/en/articles/focusvtc-efficient-and-high-performance-visual-text-compression-with-adaptive-resolution.