The Mirage of Calibrated Confidence: Trajectory-Independence of Verbalized Confidence in Vision-Language Models

Researchers found that calibrated vision-language models can still report high confidence in incorrect answers due to trajectory-independent verbalized confidence. They propose the Trajectory-Grounding Score (TGS) to address this issue, which evaluates a model's confidence based on its reasoning trajectory. This discovery highlights a blind spot in current evaluation practices for AI models.

RSS Score 0 9/17/2026, 4:00:00 AM Original Source
Save an API key to vote.