Abstract
Editable 3D models of field-grown crops support high-throughput phenotyping and in silico breeding trials, but building them from scanned point clouds requires organ-level segmentation and fitting. Procedural generators can turn an organ-level parameter set into an analysis-suitable 3D model, but obtaining that set requires hours of manual tuning per plant or segmentation models trained on species-specific labels. We present an automated pipeline that reconstructs procedural maize models from raw 3D point clouds without manual tuning or species-specific training data. A multimodal vision-language model (VLM) annotates leaf midlines in rendered orthographic views. Deterministic geometric algorithms back-project the annotations onto the point cloud, merge them into 3D leaves by cross-view consensus, and grow the midlines to full blades on an orientation-weighted surface graph. Measured organ parameters populate a plant descriptor for a Non-Uniform Rational B-Spline (NURBS)-based procedural model generator. Each leaf surface is then refined against its scan points by differentiable NURBS fitting. The pipeline reached a median whole-plant Chamfer distance of 5.4 mm on 100 genotypically diverse field-grown maize plants from the MaizeField3D dataset. The reconstructions were closer to the scans than those of an earlier semi-automated pipeline based on manual annotations. The pipeline recovered 1,017 of 1,023 (99.4%) curated reference leaves at an intersection-over-union of at least 0.5 without using those labels as input. These results show that VLM annotations become usable organ-level measurements when downstream geometric stages can correct them. This makes automated generation of editable 3D plant assets feasible at the scale of modern phenotyping experiments.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Hadadi, M., Jubery, T. Z., Krishnamurthy, A., & Ganapathysubramanian, B. (2026). A Vision-Language Model (VLM)-based Pipeline for End-to-End Procedural Modeling of Field-Grown Maize from Point Clouds. https://omanscience.com/en/articles/a-vision-language-model-vlm-based-pipeline-for-end-to-end-procedural-modeling-of-field-grown-maize-from-point-clouds
MLA 9
Hadadi, Mozhgan, et al. "A Vision-Language Model (VLM)-based Pipeline for End-to-End Procedural Modeling of Field-Grown Maize from Point Clouds." https://omanscience.com/en/articles/a-vision-language-model-vlm-based-pipeline-for-end-to-end-procedural-modeling-of-field-grown-maize-from-point-clouds.
Chicago (author–date)
Hadadi, Mozhgan, Talukder Z. Jubery, Adarsh Krishnamurthy, and Baskar Ganapathysubramanian. 2026. "A Vision-Language Model (VLM)-based Pipeline for End-to-End Procedural Modeling of Field-Grown Maize from Point Clouds." https://omanscience.com/en/articles/a-vision-language-model-vlm-based-pipeline-for-end-to-end-procedural-modeling-of-field-grown-maize-from-point-clouds.
Harvard
Hadadi, M., Jubery, T. Z., Krishnamurthy, A. and Ganapathysubramanian, B. (2026) 'A Vision-Language Model (VLM)-based Pipeline for End-to-End Procedural Modeling of Field-Grown Maize from Point Clouds', Available at: https://omanscience.com/en/articles/a-vision-language-model-vlm-based-pipeline-for-end-to-end-procedural-modeling-of-field-grown-maize-from-point-clouds.
Vancouver
Hadadi M, Jubery TZ, Krishnamurthy A, Ganapathysubramanian B. A Vision-Language Model (VLM)-based Pipeline for End-to-End Procedural Modeling of Field-Grown Maize from Point Clouds. https://omanscience.com/en/articles/a-vision-language-model-vlm-based-pipeline-for-end-to-end-procedural-modeling-of-field-grown-maize-from-point-clouds
IEEE
M. Hadadi, T. Z. Jubery, A. Krishnamurthy, and B. Ganapathysubramanian, "A Vision-Language Model (VLM)-based Pipeline for End-to-End Procedural Modeling of Field-Grown Maize from Point Clouds," https://omanscience.com/en/articles/a-vision-language-model-vlm-based-pipeline-for-end-to-end-procedural-modeling-of-field-grown-maize-from-point-clouds.