From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization
A new benchmark evaluates the safety of large language models (LLMs) in vehicle voice command authorization. The study introduces a 202-scenario benchmark and evaluates two local open-weight models and three API-based LLMs, finding that even the best-performing models produce False Executes and persistent errors. The results suggest that structured LLM decisions are insufficient as a standalone safety mechanism, and an independent enforcement layer is required to verify tool permissions and vehicle-state constraints.
Save an API key to vote.