Preprint Open access
Improving Image-Based Nutrition Estimation Through Multimodal Food-Item Verification and Recovery
Single-image nutrition estimation can fail silently when visible foods are missed. Even when a food is correctly identified, its proposed region may not support portion estimation. We propose a framework that uses multimodal large language models (MLLMs) to inventory visible foods and separately verify food identity an …