HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark

A new benchmark, HINTBench, is introduced for evaluating agent safety through non-attack intrinsic risk auditing. It includes 596 agent trajectories, with 400 synthetic risky and 136 synthetic safe trajectories, and 30 reconstructed real-world risky and safe trajectories. The benchmark supports three tasks: risk detection, risk-step localization, and intrinsic failure-type identification.

RSS Score 0 9/17/2026, 4:00:00 AM Original Source
Save an API key to vote.