نسخة أولية وصول مفتوح
Beyond Retrieval Relevance: Scene-Grounded Risk Entailment for Vision-Language Driving
Retrieval-augmented generation (RAG) gives vision--language driving systems access to external safety knowledge, yet a retrieved risk rule may be relevant without applying to the current scene. A vision--language model (VLM) receiving such knowledge must ground objects, bind entities across time, and verify relations b …