BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents

A new benchmark, Blindspot, is introduced for evaluating the safety and refusal calibration of long-horizon tool-using AI agents. It assesses agent behavior through adaptive adversarial interaction and stateful tool execution, providing a live-simulation framework for evaluating 13 LLMs and revealing substantial differences in safety-utility calibration across models.

RSS Score 0 9/16/2026, 4:00:00 AM Original Source
Save an API key to vote.