نسخة أولية وصول مفتوح
Foveated Compression: Selective High-Resolution Preservation for Token-Efficient VLMs
Visual tokens are a major source of inference cost in vision-language models, yet simple image downsampling remains a surprisingly strong compression baseline. This raises a complementary question: under a fixed token budget, where should visual fidelity be preserved? We introduce Foveated Compression, which encodes a …