Detect Before You Leap: Mirage Detection in Vision-Language Models
Researchers introduced a model-agnostic method, TC-LIA, to detect mirage reasoning in vision-language models, a failure mode where VLMs produce confident answers without relevant visual evidence. TC-LIA is an unsupervised technique that tracks question-image alignment across layers of a frozen CLIP encoder, with promising results on state-of-the-art VLMs.
Save an API key to vote.