CESBench: Benchmarking Large Language Models on Cryptographic Engineering Security for IoT Devices

CESBench is a new benchmark for evaluating the cryptographic engineering security of IoT devices using large language models (LLMs). It consists of 380 expert-written items across six sub-domains, targeting four task types: multiple-choice, judgment, scenario, and code completion. The benchmark was validated by 11 LLMs, revealing that while LLMs perform well in multiple-choice and code tasks, they struggle with justifying security verdicts. The benchmark and results are publicly available.

RSS Score 0 9/21/2026, 4:00:00 AM Original Source
Save an API key to vote.