Same Answer, Different Representations: Hidden instability in VLMs
Researchers found hidden instability in Vision Language Models (VLMs) by measuring internal representation drift, spectral sensitivity, and structural smoothness. This instability can lead to models producing unchanged outputs while undergoing significant internal changes. The study also found that larger VLMs are not necessarily more robust and that perturbations affect tasks differently, highlighting the need for more robust evaluation frameworks for VLMs.
Save an API key to vote.