BENCHCOMPASS: From Scores to Signals for Training and Harness Decisions in Payment-Domain LLMs
BENCHCOMPASS is a payment-domain benchmark that evaluates the performance of large language models in payment operations. It includes scenario-grounded tasks, LLM-based quality checks, and attack variants to assess model robustness. The benchmark shows different failure modes in 16 model variants, indicating the need for further evaluation and improvement.
Save an API key to vote.