From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization

A new benchmark evaluates the safety of large language models (LLMs) in vehicle voice command authorization. The study introduces a 202-scenario benchmark and evaluates two local open-weight models and three API-based LLMs, finding that even the best-performing models produce False Executes and persistent errors. The results suggest that structured LLM decisions are insufficient as a standalone safety mechanism, and an independent enforcement layer is required to verify tool permissions and vehicle-state constraints.

RSS Score 0 9/18/2026, 4:00:00 AM Original Source
Save an API key to vote.