The Mirage of Calibrated Confidence: Trajectory-Independence of Verbalized Confidence in Vision-Language Models
Researchers found that calibrated vision-language models can still report high confidence in incorrect answers due to trajectory-independent verbalized confidence. They propose the Trajectory-Grounding Score (TGS) to address this issue, which evaluates a model's confidence based on its reasoning trajectory. This discovery highlights a blind spot in current evaluation practices for AI models.
Save an API key to vote.