الملخص

تمت ترجمة أجزاء من هذه الصفحة آلياً وقد تحتوي على أخطاء.

Nutrition estimation is a fundamental task in consumer diet tracking, clinical dietetics, chronic disease management, sports and hospital nutrition, and broader food computing systems. The existing approaches have progressed along two largely separate axes, vision models that rely on calibrated RGB-depth captures and ingredient-aware methods that use textual cues but use limited multimodal fusion. We introduce NutriVision, an end-to-end framework that leverages visual geometry and ingredient semantics to estimate calories, mass, fat content, carbohydrates, and protein from a single RGB image and an optional ingredient list. It obtains the unavailable depth modality using DepthAnything-V3 and encodes ingredient descriptions using CLIP. It integrates three complementary mechanisms: (1) an \emph{Ingredient-Conditioned Frequency-Aligned Fusion Module (IC-FAFM)}, which uses textual guidance to reweight and align RGB-depth frequency components; (2) an \emph{Ingredient-Aware Mask-based Prediction Head (IA-MPH)}, whose gating and channel masks are conditioned on food identity; and (3) modality-specific \emph{Internal Semantic Modeling (ISM)} blocks. On the Nutrition5k dataset, NutriVision achieves a mean PMAE of $\mathbf{13.60\pm0.10\%}$, outperforming our IGSMNet implementation by $0.89$ percentage points and OmniFood8k by $2.90$ percentage points (both $p<0.001$). The module-level ablations identify the ingredient-aware prediction head as the primary architectural contributor, improving mean PMAE by $1.50\pm0.17$ percentage points ($p<0.001$). These results demonstrate that ingredient-conditioned prediction and frequency-aware RGB-depth fusion provide measurable gains for single-image nutrient estimation. More broadly, NutriVision offers a practical route toward nutrition-assessment systems that exploit geometric and semantic cues without requiring specialized depth-sensing hardware

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Kumar, A., Anand, A., Lakhchaura, C., Kumar, A., Abrol, A., Liu, T., Wang, Z., & Shah, R. R. (2026). NutriVision: الدمج والتنبؤ المشروطان بالمكونات لتقدير القيمة الغذائية من صورة واحدة. https://omanscience.com/ar/articles/nutrivision-ingredient-conditioned-fusion-and-prediction-for-single-image-food-nutrition-estimation

MLA 9

Kumar, Aman, et al. "NutriVision: الدمج والتنبؤ المشروطان بالمكونات لتقدير القيمة الغذائية من صورة واحدة." https://omanscience.com/ar/articles/nutrivision-ingredient-conditioned-fusion-and-prediction-for-single-image-food-nutrition-estimation.

شيكاغو (المؤلف–التاريخ)

Kumar, Aman, Avinash Anand, Chaitanya Lakhchaura, Ashutosh Kumar, Akshita Abrol, Timothy Liu, Zhengkui Wang, and Rajiv Ratn Shah. 2026. "NutriVision: الدمج والتنبؤ المشروطان بالمكونات لتقدير القيمة الغذائية من صورة واحدة." https://omanscience.com/ar/articles/nutrivision-ingredient-conditioned-fusion-and-prediction-for-single-image-food-nutrition-estimation.

هارفارد

Kumar, A., Anand, A., Lakhchaura, C., Kumar, A., Abrol, A., Liu, T., Wang, Z. and Shah, R. R. (2026) 'NutriVision: الدمج والتنبؤ المشروطان بالمكونات لتقدير القيمة الغذائية من صورة واحدة', Available at: https://omanscience.com/ar/articles/nutrivision-ingredient-conditioned-fusion-and-prediction-for-single-image-food-nutrition-estimation.

فانكوفر

Kumar A, Anand A, Lakhchaura C, Kumar A, Abrol A, Liu T, et al. NutriVision: الدمج والتنبؤ المشروطان بالمكونات لتقدير القيمة الغذائية من صورة واحدة. https://omanscience.com/ar/articles/nutrivision-ingredient-conditioned-fusion-and-prediction-for-single-image-food-nutrition-estimation

IEEE

A. Kumar, A. Anand, C. Lakhchaura, A. Kumar, A. Abrol, T. Liu, Z. Wang, and R. R. Shah, "NutriVision: الدمج والتنبؤ المشروطان بالمكونات لتقدير القيمة الغذائية من صورة واحدة," https://omanscience.com/ar/articles/nutrivision-ingredient-conditioned-fusion-and-prediction-for-single-image-food-nutrition-estimation.