الباحثون

Mathias Zinnen

المنشورات 1

نسخة أولية وصول مفتوح

From Pixels to Structure: Lightweight Vision-Language Models for Document OCR and Structured JSON Extraction

Uddipan Basu Bir, Vincent Christlein, Andreas Maier وآخرون · 2026 · 10.1007/978-3-032-36039-7_30

While massive, closed-source Vision-Language Models (VLMs) set strong benchmarks for document understanding, their dependence on commercial APIs limits adoption in institutional archives due to data autonomy concerns, recurring costs, and the environmental footprint of hyperscale computing. This is especially acute in …

المؤلفون المشاركون