BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents
A new benchmark, Blindspot, is introduced for evaluating the safety and refusal calibration of long-horizon tool-using AI agents. It assesses agent behavior through adaptive adversarial interaction and stateful tool execution, providing a live-simulation framework for evaluating 13 LLMs and revealing substantial differences in safety-utility calibration across models.
Save an API key to vote.