الملخص

Large language models (LLMs) are increasingly considered for safety-critical engineering, yet their reliability in regulated functional-safety workflows remains underexplored. We introduce SAFARI (Safety-Aware Functional Automotive Risk Inference), the first industrial benchmark for LLM-assisted automotive Hazard Analysis and Risk Assessment (HARA) under ISO 26262. It contains 3,000 de-identified industrial HARA cases and evaluates two coupled tasks: open-ended hazard analysis and standards-grounded risk assessment. To evaluate open-ended HARA artifacts, we propose the first reference-anchored LLM-as-a-judge protocol with high expert correlation. Experiments with nine frontier LLMs show that models often produce plausible hazard narratives but remain weak at ISO 26262 risk classification, with the best ASIL macro-F1 reaching only 0.261. Chain-of-Thought prompting provides limited benefit and often degrades categorical risk assessment. Error analysis further localizes major failures to scenario-critical context omissions during hazard generation and to controllability misjudgments during risk assessment, indicating where expert oversight should be concentrated. The dataset can be obtained from https://github.com/xixi47520-hash/HARA.

الكلمات المفتاحية

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Wu, C., Wang, Z., Zhang, H., Wang, W., & Xu, Z. (2026). SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment. https://omanscience.com/ar/articles/safari-an-industrial-benchmark-for-llm-assisted-hazard-analysis-and-risk-assessment

MLA 9

Wu, Chenxi, et al. "SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment." https://omanscience.com/ar/articles/safari-an-industrial-benchmark-for-llm-assisted-hazard-analysis-and-risk-assessment.

شيكاغو (المؤلف–التاريخ)

Wu, Chenxi, Zimu Wang, Haiyang Zhang, Wei Wang, and Zhijie Xu. 2026. "SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment." https://omanscience.com/ar/articles/safari-an-industrial-benchmark-for-llm-assisted-hazard-analysis-and-risk-assessment.

هارفارد

Wu, C., Wang, Z., Zhang, H., Wang, W. and Xu, Z. (2026) 'SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment', Available at: https://omanscience.com/ar/articles/safari-an-industrial-benchmark-for-llm-assisted-hazard-analysis-and-risk-assessment.

فانكوفر

Wu C, Wang Z, Zhang H, Wang W, Xu Z. SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment. https://omanscience.com/ar/articles/safari-an-industrial-benchmark-for-llm-assisted-hazard-analysis-and-risk-assessment

IEEE

C. Wu, Z. Wang, H. Zhang, W. Wang, and Z. Xu, "SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment," https://omanscience.com/ar/articles/safari-an-industrial-benchmark-for-llm-assisted-hazard-analysis-and-risk-assessment.