Abstract

Large language models (LLMs) are increasingly integrated into vehicle voice assistants. But linking natural-language requests to vehicle functions creates a safety-critical authorization problem. Before executing a command, the system must choose whether to execute, refuse, clarify, require confirmation, defer to manual control, trigger an emergency response, or make no tool call. To our knowledge, prior evaluations do not isolate this pre-action decision across speaker role, authentication status, vehicle state, and tool availability. We introduce a 202-scenario benchmark with Reference Decisions under a seven-class taxonomy. We evaluate two local open-weight models and three API-based LLMs using Decision Alignment and safety-specific error metrics. Alignment ranges from 40.1% for Llama 3.2 3B to 89.1% for Gemini 3.1 Pro Preview. The API-based models score between 83.2% and 89.1%, with no statistically significant differences among them. Even these models produce two to three False Executes among 161 non-execution scenarios, and persistent errors remain in confirmation and manual-control decisions. A controlled Llama 3.2 3B ablation increases alignment to 40.1% under the structured authorization policy, versus 28.2-29.2% under schema-only and generic-safety baselines, but it does not eliminate False Executes. Structured LLM decisions are therefore insufficient as a standalone safety mechanism, and deployment requires an independent enforcement layer that verifies tool permissions and vehicle-state constraints before invoking any vehicle function.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Afroze, D., Zhang, X., Tu, Y., & Hei, X. (2026). From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization. https://omanscience.com/en/articles/from-intent-to-action-benchmarking-llm-safety-in-vehicle-voice-command-authorization

MLA 9

Afroze, Diba, et al. "From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization." https://omanscience.com/en/articles/from-intent-to-action-benchmarking-llm-safety-in-vehicle-voice-command-authorization.

Chicago (author–date)

Afroze, Diba, Xingli Zhang, Yazhou Tu, and Xiali Hei. 2026. "From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization." https://omanscience.com/en/articles/from-intent-to-action-benchmarking-llm-safety-in-vehicle-voice-command-authorization.

Harvard

Afroze, D., Zhang, X., Tu, Y. and Hei, X. (2026) 'From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization', Available at: https://omanscience.com/en/articles/from-intent-to-action-benchmarking-llm-safety-in-vehicle-voice-command-authorization.

Vancouver

Afroze D, Zhang X, Tu Y, Hei X. From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization. https://omanscience.com/en/articles/from-intent-to-action-benchmarking-llm-safety-in-vehicle-voice-command-authorization

IEEE

D. Afroze, X. Zhang, Y. Tu, and X. Hei, "From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization," https://omanscience.com/en/articles/from-intent-to-action-benchmarking-llm-safety-in-vehicle-voice-command-authorization.