الباحثون

Zhengkui Wang

المنشورات 3

نسخة أولية وصول مفتوح

NutriVision: Ingredient-Conditioned Fusion and Prediction for Single-Image Food Nutrition Estimation

Nutrition estimation is a fundamental task in consumer diet tracking, clinical dietetics, chronic disease management, sports and hospital nutrition, and broader food computing systems. The existing approaches have progressed along two largely separate axes, vision models that rely on calibrated RGB-depth captures and i …

نسخة أولية وصول مفتوح

dKFD: Phase-Structured Evidence Allocation for Fixed-Budget Localized Event Understanding

Sparse video understanding often requires selecting a small set of visual evidence under a fixed frame budget. Most sparse selectors allocate this budget globally, allowing all frames to compete with one another. For temporally localized events, this can be a poor inductive bias: useful evidence is often distributed ac …

نسخة أولية وصول مفتوح

SynDORBench: Evaluating LVLM Perceptual Robustness Under Physically Constrained Visibility Conditions

Large vision-language models (LVLMs) have demonstrated remarkable performance on multimodal reasoning benchmarks, yet their perceptual reliability under physically constrained imaging conditions remains poorly understood. Existing evaluations predominantly assume ideal visual inputs and therefore fail to characterize h …

المؤلفون المشاركون