Abstract

Retrieval-augmented generation (RAG) gives vision--language driving systems access to external safety knowledge, yet a retrieved risk rule may be relevant without applying to the current scene. A vision--language model (VLM) receiving such knowledge must ground objects, bind entities across time, and verify relations before deciding how to act, leaving the support for risk conclusions implicit. We address this relevance--applicability gap with a Driving-Risk Knowledge Graph (DRKG) and Semantic Web Rule Language (SWRL) reasoning stage before VLM decision-making. Structured perception instantiates scene facts, from which SWRL rules derive events and directed risk relations when their antecedents are jointly satisfied. Recognized events, bound risk relations, and semantic descriptions of activated rules form compact evidence that conditions the VLM and diffusion planner. In matched comparisons on nuReasoning, our method improved the nuReasoning planning score (NPS) by 1.30 points and the non-at-fault collision score (NC) by 2.76 points over the relevance retrieval-based baseline. These gains indicate that scene-applicable risk evidence improves safety-weighted planning relative to semantically retrieved risk knowledge.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Liu, J., Yu, R., Peng, L., Wang, J., Zhao, C., Zhu, Z., Wang, B., Chen, G., Ye, H., Wang, H., & Li, J. (2026). Beyond Retrieval Relevance: Scene-Grounded Risk Entailment for Vision-Language Driving. https://omanscience.com/en/articles/beyond-retrieval-relevance-scene-grounded-risk-entailment-for-vision-language-driving

MLA 9

Liu, Jiaxin, et al. "Beyond Retrieval Relevance: Scene-Grounded Risk Entailment for Vision-Language Driving." https://omanscience.com/en/articles/beyond-retrieval-relevance-scene-grounded-risk-entailment-for-vision-language-driving.

Chicago (author–date)

Liu, Jiaxin, Ruilin Yu, Liang Peng, Jingkai Wang, Chengxiang Zhao, Zhenxin Zhu, Bing Wang, Guang Chen, Hangjun Ye, Hong Wang, and Jun Li. 2026. "Beyond Retrieval Relevance: Scene-Grounded Risk Entailment for Vision-Language Driving." https://omanscience.com/en/articles/beyond-retrieval-relevance-scene-grounded-risk-entailment-for-vision-language-driving.

Harvard

Liu, J., Yu, R., Peng, L., Wang, J., Zhao, C., Zhu, Z., Wang, B., Chen, G., Ye, H., Wang, H. and Li, J. (2026) 'Beyond Retrieval Relevance: Scene-Grounded Risk Entailment for Vision-Language Driving', Available at: https://omanscience.com/en/articles/beyond-retrieval-relevance-scene-grounded-risk-entailment-for-vision-language-driving.

Vancouver

Liu J, Yu R, Peng L, Wang J, Zhao C, Zhu Z, et al. Beyond Retrieval Relevance: Scene-Grounded Risk Entailment for Vision-Language Driving. https://omanscience.com/en/articles/beyond-retrieval-relevance-scene-grounded-risk-entailment-for-vision-language-driving

IEEE

J. Liu, R. Yu, L. Peng, J. Wang, C. Zhao, Z. Zhu, B. Wang, G. Chen, H. Ye, H. Wang, and J. Li, "Beyond Retrieval Relevance: Scene-Grounded Risk Entailment for Vision-Language Driving," https://omanscience.com/en/articles/beyond-retrieval-relevance-scene-grounded-risk-entailment-for-vision-language-driving.