Authors

Andreas Maier

Publications 6

Preprint Open access

From Pixels to Structure: Lightweight Vision-Language Models for Document OCR and Structured JSON Extraction

Uddipan Basu Bir, Vincent Christlein, Andreas Maier et al. · 2026 · 10.1007/978-3-032-36039-7_30

While massive, closed-source Vision-Language Models (VLMs) set strong benchmarks for document understanding, their dependence on commercial APIs limits adoption in institutional archives due to data autonomy concerns, recurring costs, and the environmental footprint of hyperscale computing. This is especially acute in …

Preprint Open access

Gradient-Based Trajectory Optimisation over Continuous Poses for Sparse-View Cone-Beam CT

Trajectory optimisation for cone-beam computed tomography (CT) determines which information sparse-view scans acquire. Fixed candidate pools prevent off-grid refinement and require new object-specific precomputation for each acquisition manifold. We make every source pose an individual continuous variable and move all …

Preprint Open access

Do Speech Representations Preserve Regional Accent Across Read and Spontaneous Speech?

Regional accent cues can be captured under matched conditions, but it remains unclear whether they persist between read and spontaneous speech. We study RVG1, with 500 German speakers from nine regions, comparing ten speech representations on regional classification and continuous geolocation under matched conditions a …

Preprint Open access

Physics-Informed but Not Physics-Consistent: Error Geometry and Subspace Projection for Neural AC Power Flow

Recent neural power-flow solvers, including emerging foundation models, achieve accurate voltage predictions, yet such accuracy does not necessarily imply physically consistent solutions. Even small complex voltage errors can yield large AC power-balance residuals. We study this accuracy-consistency gap across PIGNN-GC …

Preprint Open access

Child-Adapted Structured Phonological Representations for Interpretable Speech Sound Analysis

Structured phonological representations provide an interpretable alternative to generic speech embeddings, but existing models are largely trained on adult speech. We adapt PhonoQ-2.0 to child speech using CHILDES-Aligned data and compare three alignment-supervision conditions (Adult, Adult+Child, and Child-only) acros …

Preprint Open access

The Text Beside the Image: Detection, Utility and Leakage for Trustworthy Multimodal Medical Data and Beyond

Medical images are released with the reports that describe them, and protecting the image does not protect the report. This paper measures the text component of such releases. We measure identifier detection, downstream utility and residual identity leakage on the same documents, with the pseudonymisation policy as the …

Co-authors