Bir, U. B., Christlein, V., Maier, A., & Zinnen, M. (2026). From Pixels to Structure: Lightweight Vision-Language Models for Document OCR and Structured JSON Extraction. https://doi.org/10.1007/978-3-032-36039-7_30