نسخة أولية وصول مفتوح
From Pixels to Structure: Lightweight Vision-Language Models for Document OCR and Structured JSON Extraction
While massive, closed-source Vision-Language Models (VLMs) set strong benchmarks for document understanding, their dependence on commercial APIs limits adoption in institutional archives due to data autonomy concerns, recurring costs, and the environmental footprint of hyperscale computing. This is especially acute in …