نسخة أولية وصول مفتوح
Answer with Evidence: Consistency-Aware Grounded Visual Question Answering for Roadside Traffic Scenes
Roadside traffic reasoning requires every free-form textual claim to be backed by visual evidence. Existing grounded multimodal large language models (MLLMs) frequently exhibit say-point mismatch, in which the textual answer contradicts the bounding boxes the model localizes. Evaluation metrics that score answers and b …